Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis

#1 Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis [PDF⁵] [Copy] [Kimi²] [REL]

Authors: Sho Inoue, Kun Zhou, Shuai Wang, Haizhou Li

We investigate hierarchical emotion distribution (ED) for achieving multi-level quantitative control of emotion rendering in text-to-speech synthesis (TTS). We introduce a novel multi-step hierarchical ED prediction module that quantifies emotion variance at the utterance, word, and phoneme levels. By predicting emotion variance in a multi-step manner, we leverage global emotional context to refine local emotional variations, thereby capturing the intrinsic hierarchical structure of speech emotion. Our approach is validated through its integration into a variance adaptor and an external module design compatible with various TTS systems. Both objective and subjective evaluations demonstrate that the proposed framework significantly enhances emotional expressiveness and enables precise control of emotion rendering across multiple speech granularities.

Subjects: Sound , Audio and Speech Processing

Publish: 2025-07-07 01:25:51 UTC

2507.04598

#1 Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis [PDF5] [Copy] [Kimi2] [REL]

#1 Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis [PDF⁵] [Copy] [Kimi²] [REL]