Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation

#1 Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation [PDF⁷] [Copy] [Kimi⁹] [REL]

Authors: Xinhan Di, Zihao Chen, Yunming Liang, Junjie Zheng, Yihua Wang, Chaofan Ding

Large-scale text-to-speech (TTS) models have made significant progress recently.However, they still fall short in the generation of Chinese dialectal speech. Toaddress this, we propose Bailing-TTS, a family of large-scale TTS models capable of generating high-quality Chinese dialectal speech. Bailing-TTS serves as a foundation model for Chinese dialectal speech generation. First, continual semi-supervised learning is proposed to facilitate the alignment of text tokens and speech tokens. Second, the Chinese dialectal representation learning is developed using a specific transformer architecture and multi-stage training processes. With the proposed design of novel network architecture and corresponding strategy, Bailing-TTS is able to generate Chinese dialectal speech from text effectively and efficiently. Experiments demonstrate that Bailing-TTS generates Chinese dialectal speech towards human-like spontaneous representation. Readers are encouraged to listen to demos at \url{https://c9412600.github.io/bltts_tech_report/index.html}.

Subjects: Computation and Language , Sound , Audio and Speech Processing

Publish: 2024-08-01 04:57:31 UTC

2408.00284

#1 Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation [PDF7] [Copy] [Kimi9] [REL]

#1 Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation [PDF⁷] [Copy] [Kimi⁹] [REL]