li26l@interspeech_2026@ISCA

Total: 1

#1 HFMSE: Harmonic-Guided Speech Enhancement with Flow Matching [PDF] [Copy] [Kimi] [REL]

Authors: Jizhen Li, Weiping Tu, Yuhong Yang, Xinhong Li

Generative speech enhancement methods have shown impressive performance by directly modeling clean speech distribution, yet their effectiveness critically depends on the reliability of conditional information. Mainstream conditioning strategy face two primary challenges: the shallow features extracted from noisy speech often fails to capture structured acoustic features such as harmonic structures, and extracting accurate and reliable semantic information from noisy signals poses a challenge comparable to the enhancement task itself. To address these limitations, we propose HFMSE, a Harmonic-guided Flow Matching method for Speech Enhancement. Specifically, we design an efficient harmonic encoder that extracts harmonic structural features from noisy speech through a two-step process of fundamental frequency localization and harmonic mask generation. These features serve as strong conditioning guidance and are integrated into the flow matching generation process to improve the structural integrity and perceptual naturalness of the output speech. Extensive experiments on DNS Challenge 2020 show that HFMSE achieves state-of-the-art performance, demonstrating its superior capability and robustness. Code is available at https://github.com/xxnhq/HFSE/.

Subject: INTERSPEECH.2026 - Speech Synthesis