wang26ea@interspeech_2026@ISCA

Total: 1

#1 Not Flat, But Dissociated: Prosodic and Segmental Divergence in Neural TTS [PDF] [Copy] [Kimi] [REL]

Authors: Rong Wang, Kun Sun, Harald Baayen

Standard TTS metrics such as MOS and mel-cepstral distortion provide global scores but do not locate where synthesis diverges from natural speech. To examine this, four systems (Tacotron2-DDC, FastSpeech2, Glow-TTS, MixerTTS) are analysed across 13,100 matched LJ-TTS utterances. At the prosodic level, global F0 variability is compressed (d = −0.55) while local pitch reversals increase (d = +0.82), suggesting a cross-timescale dissociation, rather than monotone intonation. Timing metrics show no reliable deviation, localizing the deficit to F0 coordination. At the segmental level, vowel spaces shrink to 9–30% of the human baseline, directional formant biases indicate articulatory undershoot, and locus analysis confirms reduced place-conditioned coarticulation most consistently for alveolars. Deviations at these two levels are largely uncorrelated, suggesting prosodic organisation and segmental precision are distinct dimensions of synthesis quality conflated by standard metrics.

Subject: INTERSPEECH.2026 - Speech Synthesis