Emotional TTS
Research

Evaluation

Evaluation methodology, datasets, and metrics used to assess the Emotional TTS system.

Evaluation Datasets

CV3_Eval

Zero-shot

Evaluation with unseen speakers not included in training data. Tests the model's ability to generalize to new speaker identities using only a short (~5s) reference clip.

EmoVoiceDB

Zero-shot

A dedicated emotional voice database for evaluation. Tests both reference-based and vector-based emotion control with unseen speakers across multiple emotions.

ESD (Emotional Speech Dataset)

Seen Speakers

Evaluation with speakers included in training data. Tests emotion expressiveness and quality when speaker adaptation is not the primary challenge.

Evaluation Methods

Reference-Based

Emotion is transferred from a reference audio sample. Evaluates the quality of emotion cloning while maintaining speaker identity.

Vector-Based

Emotion is controlled via predefined embeddings. Only a neutral speaker reference is provided. Evaluates controllability and emotion consistency.

Intensity Control

Tests the system's ability to modulate emotion intensity using the emo_alpha parameter (0.33, 0.66, 1.0).

Emotions Tested

HappyAngrySadAfraidSurprisedCalmDisgusted

Evaluation Criteria

  • Speaker Similarity: How well the generated speech preserves the identity of the reference speaker
  • Emotion Expressiveness: How clearly the target emotion is expressed in the output
  • Naturalness: Overall quality and naturalness of the synthesized speech
  • Intelligibility: Clarity and understandability of the generated speech

Results Summary

DatasetSettingMethodsSamples
CV3_EvalZero-shotReference-based10
EmoVoiceDBZero-shotReference + Vector30+
ESDSeen speakersReference + Vector + Intensity40+

Zero-Shot Performance Comparison

LibriSpeech test-clean

ModelSS↑WER(%)↓SMOS↑PMOS↑QMOS↑
Ground Truth0.8333.4054.02±0.223.85±0.264.23±0.12
MaskGCT0.7907.7594.12±0.093.98±0.114.19±0.19
F5-TTS0.8218.0444.08±0.213.73±0.274.12±0.13
CosyVoice 20.8435.9994.02±0.224.04±0.284.17±0.25
SparkTTS0.7568.8434.06±0.203.94±0.214.15±0.16
IndexTTS0.8193.4364.23±0.144.02±0.184.29±0.22
IndexTTS 20.8703.1154.44±0.124.12±0.174.29±0.14

Emotional Test Set

Emotional Test Set

ModelSS↑WER(%)↓ES↑SMOS↑EMOS↑PMOS↑QMOS↑
MaskGCT0.8104.0590.8413.42±0.363.37±0.423.04±0.403.39±0.37
F5-TTS0.7733.0530.7573.37±0.313.16±0.323.13±0.403.36±0.29
CosyVoice 20.8031.8310.8023.13±0.323.09±0.332.98±0.353.28±0.22
SparkTTS0.6732.2990.8323.01±0.263.16±0.343.21±0.283.04±0.18
IndexTTS0.6491.1360.6603.17±0.392.74±0.363.15±0.363.56±0.27
IndexTTS 20.8361.8830.8874.24±0.194.22±0.124.08±0.204.18±0.10

Natural Language Emotion Control

Natural Language Emotion Control

ModelSMOS↑EMOS↑PMOS↑QMOS↑
CosyVoice 22.973±0.263.339±0.303.679±0.193.429±0.24
IndexTTS 23.875±0.213.786±0.244.143±0.134.071±0.15

Emotion Transfer (ESD Dataset)

Emotion Prediction Accuracy

HappySurprisedSadAfraidAngry
IndexTTS0.9000.9330.7830.6941.000

Speaker Similarity

HappySurprisedSadAfraidAngry
IndexTTS0.9540.9440.9350.9010.970

Emotion Vector Similarity

HappySurprisedSadAfraidAngry
IndexTTS0.8780.8290.8410.7800.969

MCD Score (dB) — Lower is better

HappySurprisedSadAfraidAngry
IndexTTS4.65.45.346.34.6

Cross-Model Comparison (CV3-Eval)

Emotion Prediction Accuracy

HappySadAngry
CosyVoice 364.3350.167.45
ChatTTS65.8543.5358.97
IndexTTS68.2953.8594.87

Emotion Vector Similarity

HappySadAngry
CosyVoice 381.3476.381.5
ChatTTS82.1878.9382.48
IndexTTS86.8884.7790.72

Speaker Similarity

HappySadAngry
CosyVoice 393.5594.393.8
ChatTTS92.5593.8294.09
IndexTTS93.3894.5393.56

EmoVoiceDB Evaluation

Emotion Prediction Accuracy

HappySurprisedSadAfraidAngry
IndexTTS0.8490.4820.5630.4230.808

Speaker Similarity

HappySurprisedSadAfraidAngry
IndexTTS0.9390.8530.8640.8890.917