Total: 1
Deep learning-based speech enhancement (SE) models are typically trained on synthetic pairs of clean speech and synthesized noisy speech, which often generalize poorly to real-world acoustic environments. In practical target environments, however, paired clean target speech aligned with noisy recordings is usually unavailable, making it difficult to directly use real data for supervised SE training. To address these limitations, we propose SwitchSE — a target-domain clean-free fine-tuning framework with a switch-controlled mechanism that leverages transcription-noisy speech pairs to upgrade a pre-trained SE model in realistic environments. Experiments show that with only 2.9 h of CHiME-3 real noisy speech, SwitchSE substantially improves target-domain performance while retaining strong performance on the original synthetic dataset, providing a data-efficient framework to bridge the gap between synthesized training data and real-world application demands.