yuan26b@interspeech_2026@ISCA

Total: 1

#1 DelayGSE: A Generative Speech Enhancement Framework with Delayed Text-Aware Conditioning [PDF] [Copy] [Kimi] [REL]

Authors: Xin Yuan, Junling Lv, Zezhou Xu, Xingjun Tan, Liangliang Li, Yanqiang Lei

Recent generative speech enhancement methods based on language and diffusion models achieve strong perceptual quality but are more susceptible than discriminative approaches to speech-like hallucinations under low SNRs and transient noise. We propose DelayGSE, a text-aware generative speech enhancement framework built on a multi-codebook language model for denoising, dereverberation, and audio super-resolution. DelayGSE conditions on noisy-speech STFT features and Whisper encoder representations, and models multiple discrete codebooks in a delayed manner to stabilize generation. A text-aware mechanism suppresses hallucinations, while an importance-aware codebook weighting strategy balances perceptual fidelity and semantic consistency. Experiments demonstrate state-of-the-art performance, with ablations showing effective hallucination suppression and a 15.8% relative word error rate reduction. Audio samples are available at https://delaygse.github.io/.

Subject: INTERSPEECH.2026 - Speech Processing