shang26@interspeech_2026@ISCA

Total: 1

#1 Seed-Enh: Generative Speech Enhancement in Decoupled Semantic and Timbre Spaces [PDF] [Copy] [Kimi] [REL]

Authors: Zengqiang Shang, Biao Liu, Yu Zhao, Pengyuan Zhang

Most existing speech enhancement operate directly in the acoustic space, where noise and speech components are inherently entangled, leading to artifacts such as high-frequency attenuation and background holes. In this paper, we propose Seed-Enh, a generative speech enhancement framework that performs enhancement in decoupled semantic and timbre spaces. Our framework processes noisy speech in three stages: (1) semantic space processing using a frozen Whisper encoder to extract noise-robust semantic representations, (2) timbre space processing using CAM++ embeddings and context learning, and (3) space fusion via flow matching. Experiments demonstrate that Seed-Enh achieves the highest overall quality scores compared to state-of-the-art baselines, effectively suppressing noise while restoring harmonic structures and high-frequency details. Beyond speech enhancement, Seed-Enh also enables zero-shot voice conversion, significantly outperforming Seed-VC when handling noisy inputs. The demo and pretrained model are available at https://github.com/shangqwe123/Seed-Enh.

Subject: INTERSPEECH.2026 - Speech Synthesis