Total: 1
Existing text-to-speech systems predominantly focus on single-sentence synthesis and lack adequate contextual modeling and fine-grained control for coherent multicast audiobooks. To address this, we propose a context-aware, emotion-controllable speech synthesis framework with three innovations: a context mechanism for consistency, a disentanglement paradigm to decouple style from prompts, and self-distillation to boost expressiveness. Experiments show consistent gains: chapter-level generation achieves 4.25 M-MOS(14% relative gain over the strongest baseline), dialog reaches 4.11 S-MOS, and emotion control improves by 31 percentage points in high-intensity discrimination. Ablation studies validate our methods. Demo: https://semisemi-ux.github.io/.