Total: 1
Achieving fine-grained emphasis control in speech synthesis remains challenging due to data scarcity and the inherent complexity of prosody. To address this issue, we extend F5-TTS with an additional Emphasis Encoder and propose a three-stage optimization framework that progressively enhances emphasis controllability. The framework begins with Supervised Fine-Tuning on manually annotated data, followed by Direct Preference Optimization, where preference pairs are constructed by ranking SFT-generated samples using the Wavelet Prosody Toolkit (WPT). In the final stage, we adapt Flow-CPS to perform online reinforcement learning (RL) for flow-matching models. Using WPT as a reward model, this stage refines the flow trajectories through the estimation of the relative advantage of the group. Experimental results demonstrate that the proposed pipeline substantially improves emphasis intensity and controllability while preserving natural prosody. Audio samples are available at https://thuhcsi.github.io/interspeech2026-F5Emphasis.