Total: 1
Most existing neural speech codecs are optimized for fixed sampling rates and lack multi-rate adaptability, leading to significant performance degradation at mismatched rates. To address this, we propose a multi-sampling-rate neural speech codec supporting 16/24/48 kHz audio in a unified model. The architecture consists of a shared encoder-decoder backbone and a universal Residual Vector Quantization (RVQ) module, integrated with two lightweight sampling-rate-aware components: an adapter for learnable waveform conversion to a unified internal temporal grid, and a transformation modulator for intermediate feature calibration to ensure cross-sampling-rate representation consistency for shared quantization. We further adopt a three-stage progressive training strategy to stabilize optimization. Experiments show that our unified codec matches the performance of sampling-rate-specific models while simplifying deployment by eliminating the need for multiple sampling-rate-specific instances.