jin26b@interspeech_2026@ISCA

Total: 1

#1 Learning to Rescale: On-the-Fly Sequence Length Adaptation in Non-Autoregressive Speech Synthesis [PDF] [Copy] [Kimi] [REL]

Authors: Jiawei Jin, Ren Wang, Zhiyu Cui, Shun Lei, Yixuan Zhou, Zhiyong Wu

Non-autoregressive (NAR) text-to-speech (TTS) models excel in parallel inference and style consistency. However, their reliance on accurate character-level or global duration predictions remains a critical bottleneck, limiting both synthesis fidelity and naturalness. This paper proposes ElasticDLM (Elastic-length Diffusion Language Model), a novel variable-length NAR TTS model that eliminates the strong dependency on accurate duration predictions. We introduce two functional tokens, along with a Differentiated Length-Scaling Training Scheme (DLTS) and a Hierarchical Confidence-Guided Inference (HCGI) strategy tailored for speech characteristics, enabling the model to learn length-scaling features and automatically adjust sequence length during inference. Experiments show that ElasticDLM produces high-fidelity, natural speech from arbitrary-length inputs, demonstrating a more flexible and simplified approach to high-quality speech synthesis. Audio samples are available at https://thuhcsi.github.io/ElasticDLM/.

Subject: INTERSPEECH.2026 - Speech Synthesis