Total: 1
Standard TTS metrics such as MOS and mel-cepstral distortion provide global scores but do not locate where synthesis diverges from natural speech. To examine this, four systems (Tacotron2-DDC, FastSpeech2, Glow-TTS, MixerTTS) are analysed across 13,100 matched LJ-TTS utterances. At the prosodic level, global F0 variability is compressed (d = −0.55) while local pitch reversals increase (d = +0.82), suggesting a cross-timescale dissociation, rather than monotone intonation. Timing metrics show no reliable deviation, localizing the deficit to F0 coordination. At the segmental level, vowel spaces shrink to 9–30% of the human baseline, directional formant biases indicate articulatory undershoot, and locus analysis confirms reduced place-conditioned coarticulation most consistently for alveolars. Deviations at these two levels are largely uncorrelated, suggesting prosodic organisation and segmental precision are distinct dimensions of synthesis quality conflated by standard metrics.