liu26p@interspeech_2026@ISCA

Total: 1

#1 audiobook-cc: Controllable Long-context Speech Generation for Multicast Audiobook [PDF] [Copy] [Kimi] [REL]

Authors: Min Liu, JingJing Yin, Xiang Zhang, JianHao Ye, Siyu Hao, Siwei Xia, Hongbin Zhou

Existing text-to-speech systems predominantly focus on single-sentence synthesis and lack adequate contextual modeling and fine-grained control for coherent multicast audiobooks. To address this, we propose a context-aware, emotion-controllable speech synthesis framework with three innovations: a context mechanism for consistency, a disentanglement paradigm to decouple style from prompts, and self-distillation to boost expressiveness. Experiments show consistent gains: chapter-level generation achieves 4.25 M-MOS(14% relative gain over the strongest baseline), dialog reaches 4.11 S-MOS, and emotion control improves by 31 percentage points in high-intensity discrimination. Ablation studies validate our methods. Demo: https://semisemi-ux.github.io/.

Subject: INTERSPEECH.2026 - Speech Synthesis