yang26l@interspeech_2026@ISCA

Total: 1

#1 CraftTTS: Fine-Grained Prosody Control for Text-to-Speech [PDF] [Copy] [Kimi] [REL]

Authors: Wenbing Yang, Qihang Lu, Bingsong Bai, Zihan Sun, Yueran Hou, Peilei Jia, Yingming Gao, Ya Li, Jun Gao

While zero-shot text-to-speech models perform well in global voice cloning, they struggle with fine-grained prosodic control, as strict word-level intensity and tempo manipulation can disrupt acoustic priors and introduce artifacts. We propose CraftTTS, a three-stage framework enabling stable word-level control without sacrificing fluency. First, a compute-driven zero-shot pipeline automatically constructs large-scale preference pairs without manual annotation. Second, joint supervised fine-tuning and direct preference optimization improve sensitivity to local prosodic tags. Third, group relative policy optimization with a multi-dimensional prosodic reward balances local controllability and global naturalness by regulating intelligibility, intensity contrast, and rhythm. Experiments demonstrate that CraftTTS achieves state-of-the-art fine-grained expressiveness while preserving zero-shot generation capability.

Subject: INTERSPEECH.2026 - Speech Synthesis