huang26l@interspeech_2026@ISCA

Total: 1

#1 Unified Neural Speech Coding for Multiple Sampling Rates [PDF] [Copy] [Kimi] [REL]

Authors: Jiankai Huang, Junteng Zhang, Lizhong Wang, Liang Wen, Ming Lu, Zhan Ma

Most existing neural speech codecs are optimized for fixed sampling rates and lack multi-rate adaptability, leading to significant performance degradation at mismatched rates. To address this, we propose a multi-sampling-rate neural speech codec supporting 16/24/48 kHz audio in a unified model. The architecture consists of a shared encoder-decoder backbone and a universal Residual Vector Quantization (RVQ) module, integrated with two lightweight sampling-rate-aware components: an adapter for learnable waveform conversion to a unified internal temporal grid, and a transformation modulator for intermediate feature calibration to ensure cross-sampling-rate representation consistency for shared quantization. We further adopt a three-stage progressive training strategy to stabilize optimization. Experiments show that our unified codec matches the performance of sampling-rate-specific models while simplifying deployment by eliminating the need for multiple sampling-rate-specific instances.

Subject: INTERSPEECH.2026 - Speech Processing