Total: 1
Non-autoregressive (NAR) text-to-speech (TTS) models excel in parallel inference and style consistency. However, their reliance on accurate character-level or global duration predictions remains a critical bottleneck, limiting both synthesis fidelity and naturalness. This paper proposes ElasticDLM (Elastic-length Diffusion Language Model), a novel variable-length NAR TTS model that eliminates the strong dependency on accurate duration predictions. We introduce two functional tokens, along with a Differentiated Length-Scaling Training Scheme (DLTS) and a Hierarchical Confidence-Guided Inference (HCGI) strategy tailored for speech characteristics, enabling the model to learn length-scaling features and automatically adjust sequence length during inference. Experiments show that ElasticDLM produces high-fidelity, natural speech from arbitrary-length inputs, demonstrating a more flexible and simplified approach to high-quality speech synthesis. Audio samples are available at https://thuhcsi.github.io/ElasticDLM/.