Total: 1
Generative speech enhancement has shown significant advancements in improving speech quality in noisy environments. However, its iterative nature suffers from high inference computational complexity. In this paper, we propose EffDiffSE+, an efficient single-iteration diffusion-based speech enhancement model based on a condition DNN, a bridge DNN, and our novel learnable Schrödinger bridge (SB). Our contributions are threefold. First, we propose a single-iteration SB with Gaussian distribution initialization for the reverse process. Second, an auxiliary network is proposed to provide learnable adaptivity to the bridge DNN initial state estimation. Third, topology improvements are presented, constituting our EffDiffSE+ to consistently improve model performance. The proposed EffDiffSE+ model excels top open-source time- and frequency-domain diffusion baseline methods in PESQ, POLQA, NISQA, UTMOS, ESTOI, LPS, SBScore, SpkSim, and subjective MOS, clearly achieving an overall top rank.