Evaluation
Evaluation methodology, datasets, and metrics used to assess the Emotional TTS system.
Evaluation Datasets
CV3_Eval
Zero-shotEvaluation with unseen speakers not included in training data. Tests the model's ability to generalize to new speaker identities using only a short (~5s) reference clip.
EmoVoiceDB
Zero-shotA dedicated emotional voice database for evaluation. Tests both reference-based and vector-based emotion control with unseen speakers across multiple emotions.
ESD (Emotional Speech Dataset)
Seen SpeakersEvaluation with speakers included in training data. Tests emotion expressiveness and quality when speaker adaptation is not the primary challenge.
Evaluation Methods
Reference-Based
Emotion is transferred from a reference audio sample. Evaluates the quality of emotion cloning while maintaining speaker identity.
Vector-Based
Emotion is controlled via predefined embeddings. Only a neutral speaker reference is provided. Evaluates controllability and emotion consistency.
Intensity Control
Tests the system's ability to modulate emotion intensity using the emo_alpha parameter (0.33, 0.66, 1.0).
Emotions Tested
Evaluation Criteria
- Speaker Similarity: How well the generated speech preserves the identity of the reference speaker
- Emotion Expressiveness: How clearly the target emotion is expressed in the output
- Naturalness: Overall quality and naturalness of the synthesized speech
- Intelligibility: Clarity and understandability of the generated speech
Results Summary
| Dataset | Setting | Methods | Samples |
|---|---|---|---|
| CV3_Eval | Zero-shot | Reference-based | 10 |
| EmoVoiceDB | Zero-shot | Reference + Vector | 30+ |
| ESD | Seen speakers | Reference + Vector + Intensity | 40+ |
Zero-Shot Performance Comparison
LibriSpeech test-clean
| Model | SS↑ | WER(%)↓ | SMOS↑ | PMOS↑ | QMOS↑ |
|---|---|---|---|---|---|
| Ground Truth | 0.833 | 3.405 | 4.02±0.22 | 3.85±0.26 | 4.23±0.12 |
| MaskGCT | 0.790 | 7.759 | 4.12±0.09 | 3.98±0.11 | 4.19±0.19 |
| F5-TTS | 0.821 | 8.044 | 4.08±0.21 | 3.73±0.27 | 4.12±0.13 |
| CosyVoice 2 | 0.843 | 5.999 | 4.02±0.22 | 4.04±0.28 | 4.17±0.25 |
| SparkTTS | 0.756 | 8.843 | 4.06±0.20 | 3.94±0.21 | 4.15±0.16 |
| IndexTTS | 0.819 | 3.436 | 4.23±0.14 | 4.02±0.18 | 4.29±0.22 |
| IndexTTS 2 | 0.870 | 3.115 | 4.44±0.12 | 4.12±0.17 | 4.29±0.14 |
Emotional Test Set
Emotional Test Set
| Model | SS↑ | WER(%)↓ | ES↑ | SMOS↑ | EMOS↑ | PMOS↑ | QMOS↑ |
|---|---|---|---|---|---|---|---|
| MaskGCT | 0.810 | 4.059 | 0.841 | 3.42±0.36 | 3.37±0.42 | 3.04±0.40 | 3.39±0.37 |
| F5-TTS | 0.773 | 3.053 | 0.757 | 3.37±0.31 | 3.16±0.32 | 3.13±0.40 | 3.36±0.29 |
| CosyVoice 2 | 0.803 | 1.831 | 0.802 | 3.13±0.32 | 3.09±0.33 | 2.98±0.35 | 3.28±0.22 |
| SparkTTS | 0.673 | 2.299 | 0.832 | 3.01±0.26 | 3.16±0.34 | 3.21±0.28 | 3.04±0.18 |
| IndexTTS | 0.649 | 1.136 | 0.660 | 3.17±0.39 | 2.74±0.36 | 3.15±0.36 | 3.56±0.27 |
| IndexTTS 2 | 0.836 | 1.883 | 0.887 | 4.24±0.19 | 4.22±0.12 | 4.08±0.20 | 4.18±0.10 |
Natural Language Emotion Control
Natural Language Emotion Control
| Model | SMOS↑ | EMOS↑ | PMOS↑ | QMOS↑ |
|---|---|---|---|---|
| CosyVoice 2 | 2.973±0.26 | 3.339±0.30 | 3.679±0.19 | 3.429±0.24 |
| IndexTTS 2 | 3.875±0.21 | 3.786±0.24 | 4.143±0.13 | 4.071±0.15 |
Emotion Transfer (ESD Dataset)
Emotion Prediction Accuracy
| Happy | Surprised | Sad | Afraid | Angry | |
|---|---|---|---|---|---|
| IndexTTS | 0.900 | 0.933 | 0.783 | 0.694 | 1.000 |
Speaker Similarity
| Happy | Surprised | Sad | Afraid | Angry | |
|---|---|---|---|---|---|
| IndexTTS | 0.954 | 0.944 | 0.935 | 0.901 | 0.970 |
Emotion Vector Similarity
| Happy | Surprised | Sad | Afraid | Angry | |
|---|---|---|---|---|---|
| IndexTTS | 0.878 | 0.829 | 0.841 | 0.780 | 0.969 |
MCD Score (dB) — Lower is better
| Happy | Surprised | Sad | Afraid | Angry | |
|---|---|---|---|---|---|
| IndexTTS | 4.6 | 5.4 | 5.34 | 6.3 | 4.6 |
Cross-Model Comparison (CV3-Eval)
Emotion Prediction Accuracy
| Happy | Sad | Angry | |
|---|---|---|---|
| CosyVoice 3 | 64.33 | 50.1 | 67.45 |
| ChatTTS | 65.85 | 43.53 | 58.97 |
| IndexTTS | 68.29 | 53.85 | 94.87 |
Emotion Vector Similarity
| Happy | Sad | Angry | |
|---|---|---|---|
| CosyVoice 3 | 81.34 | 76.3 | 81.5 |
| ChatTTS | 82.18 | 78.93 | 82.48 |
| IndexTTS | 86.88 | 84.77 | 90.72 |
Speaker Similarity
| Happy | Sad | Angry | |
|---|---|---|---|
| CosyVoice 3 | 93.55 | 94.3 | 93.8 |
| ChatTTS | 92.55 | 93.82 | 94.09 |
| IndexTTS | 93.38 | 94.53 | 93.56 |
EmoVoiceDB Evaluation
Emotion Prediction Accuracy
| Happy | Surprised | Sad | Afraid | Angry | |
|---|---|---|---|---|---|
| IndexTTS | 0.849 | 0.482 | 0.563 | 0.423 | 0.808 |
Speaker Similarity
| Happy | Surprised | Sad | Afraid | Angry | |
|---|---|---|---|---|---|
| IndexTTS | 0.939 | 0.853 | 0.864 | 0.889 | 0.917 |