INTERSPEECH.2026 - Speech Processing

| Total: 143

#1 VoCodec: A Low-bitrate Streamable Neural Speech Codec with Voicing-driven Quantization [PDF] [Copy] [Kimi] [REL]

Authors: Xiao-Hang Jiang, Yang Ai, Rui-Chen Zheng, Lirong Dai, Zhen-Hua Ling, Ji Wu

Neural speech codecs are key to speech transmission and storage, but most use uniform quantization across frames, allocating the same bitrate regardless of content and wasting bits. We propose VoCodec, a low-bitrate streamable neural speech codec with voicing-driven quantization that assigns higher bitrate to voiced frames and lower bitrate to unvoiced frames according to perceptual sensitivity. VoCodec embeds a voicing detector in a fully causal encoder-quantizer-decoder neural coding framework, using residual scalar-vector quantization for voiced frames and simple scalar quantization for unvoiced ones. Experiments show that on the LibriTTS dataset at a 16 kHz sampling rate, VoCodec outperforms baseline neural speech codecs even at a bitrate as low as 1.1 kbps. Our further experiments also confirm that introducing voicing-driven quantization can effectively reduce the bitrate by approximately 27% compared with uniform quantization strategy.

Subject: INTERSPEECH.2026 - Speech Processing


#2 Noisy Environment Adaptation of Neural Speech Codec via Focal Mask and Noise Feature Separation [PDF] [Copy] [Kimi] [REL]

Authors: Shaokai Li, Weiping Tu, Yuhong Yang

Neural speech codec has attracted extensive attention for high-quality reconstruction at low-bitrate. However, real-world noise severely degrades its performance and hinders high-quality clean speech reconstruction. To tackle this problem, we propose FocalSE, a novel speech enhancement method that performs feature denoising, noise feature separation and noise recognition in the continuous embedding space of neural speech codecs. Specifically, we develop focal modulation-based compression and decompression to capture global context and local mutual information, and generate focal masks to recover clean feature embeddings. We then separate noise embeddings from noisy embeddings to improve denoising performance. Finally, we use ResNet1D-18 to recognize noise categories for better separation effectiveness. Extensive experiments on two standard datasets, LibriTTS and ESC50, demonstrate that our method outperforms state-of-the-art approaches under low-bitrate and low-SNR conditions.

Subject: INTERSPEECH.2026 - Speech Processing


#3 Pitch-Injected Residual Adapter for Tonal Language in Neural Audio Codec [PDF] [Copy] [Kimi] [REL]

Authors: Jie-Shiang Yang, Ya-Tse Wu, Chi-Chun Lee

Neural Audio Codecs (NACs) achieve high-fidelity reconstruction at low bitrates but remain insensitive to fundamental frequency (F0) distortion—a critical limitation for tonal languages, covering 40%+ of the world's languages, where F0 contour marks lexical distinctions. We propose the Pitch-Injected Residual Adapter (PIRA), a lightweight plug-and-play module that restores tonal information in frozen NACs through explicit F0 and voiced/unvoiced side-information injection. PIRA processes quantized F0 through dilated convolutions to capture tone sandhi dependencies, supervised by a CREPE embedding loss for pitch-specific gradients. A confidence network gates injection per-frame, suppressing it for checked tones and unvoiced segments. With 1.25M–1.65M trainable parameters, PIRA reduces codec-induced Tone Error Rate (dTER) by 35.7% on average (macro-averaged across five codecs, each averaged over three tonal languages), while preserving English quality and adding ≤0.4 kbps overhead.

Subject: INTERSPEECH.2026 - Speech Processing


#4 BridgeCodec: Mamba Enhanced Neural Audio Codec with Schrödinger Bridge at Low Bitrate [PDF] [Copy] [Kimi] [REL]

Authors: Zijian Lin, Jing Yang, Jinghao Luo, Zhuo Wang, Fan Fan, Zhiyong Wu

Neural audio codecs achieve remarkable performance but operate as closed ecosystems with rigidly co-trained encoder-decoder pairs, severely hindering interoperability. To resolve this, we propose BridgeCodec, a versatile framework that decouples these pairs by translating latent representations between mismatched, frozen endpoints. We formulate this cross-codec translation as a Schrödinger Bridge optimal transport problem, enabling robust mapping between disparate distributions without predefined noise priors. To efficiently capture long-range speech temporal dependencies, we integrate a Mamba-enhanced U-Net architecture. We validate BridgeCodec on an extreme deployment scenario: mapping a lightweight 8 kHz source encoder directly to a 48 kHz target decoder at an ultra-low bitrate of 1 kbps. Experimental results demonstrate that BridgeCodec achieves superior wideband reconstruction and high perceptual quality, successfully bridging heavily bottlenecked, heterogeneous codecs.

Subject: INTERSPEECH.2026 - Speech Processing


#5 AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation [PDF] [Copy] [Kimi] [REL]

Authors: Dongmei Wang, Xiaohang Sun, Yang Liu, Fanjie Kong, Abhishek Yanamandra, Abhinav Jain, Daniel Tompkins, Woohyun Kang, Najmeh Sadoughi, Sunil Hadap, Xiang Hao, Zhu Liu, Caren Chen

We propose AugCodec, a low-bitrate disentangled neural speech codec that leverages data augmentation to decompose speech into three distinct components: semantic, speaker, and prosody tokens. Specifically, we employ tailored augmentation strategies to transform speech into distinct variants, each serving as input for extracting tokens that preserve the target attribute while suppressing others. This disentanglement strategy enables substantial reduction in token rate. Furthermore, we introduce an augmentation loss that aligns semantic encoder outputs between source and voice-converted speech, encouraging speaker-agnostic embeddings while mitigating the acoustic mismatch induced by voice conversion. Experiments on LibriSpeech test-clean demonstrate that AugCodec significantly outperforms state-of-the-art methods in both reconstruction quality and disentanglement, while operating at only 12.5Hz with three token streams.

Subject: INTERSPEECH.2026 - Speech Processing


#6 Unified Neural Speech Coding for Multiple Sampling Rates [PDF] [Copy] [Kimi] [REL]

Authors: Jiankai Huang, Junteng Zhang, Lizhong Wang, Liang Wen, Ming Lu, Zhan Ma

Most existing neural speech codecs are optimized for fixed sampling rates and lack multi-rate adaptability, leading to significant performance degradation at mismatched rates. To address this, we propose a multi-sampling-rate neural speech codec supporting 16/24/48 kHz audio in a unified model. The architecture consists of a shared encoder-decoder backbone and a universal Residual Vector Quantization (RVQ) module, integrated with two lightweight sampling-rate-aware components: an adapter for learnable waveform conversion to a unified internal temporal grid, and a transformation modulator for intermediate feature calibration to ensure cross-sampling-rate representation consistency for shared quantization. We further adopt a three-stage progressive training strategy to stabilize optimization. Experiments show that our unified codec matches the performance of sampling-rate-specific models while simplifying deployment by eliminating the need for multiple sampling-rate-specific instances.

Subject: INTERSPEECH.2026 - Speech Processing


#7 ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding [PDF] [Copy] [Kimi] [REL]

Authors: Chengbin Liang, Wenqi Guo, Hao Cao, Zhijin Qin

Neural speech codecs enable low-bitrate speech communication, yet at ultra-low bitrates (< 1000 bps) preserving perceptual quality and intelligibility is challenging. Existing designs often prioritize acoustic details, leaving limited capacity for the core linguistic message under tight bitrate constraints. To address this, we propose ContextCodec, a codec that transmits content-focused context features to explicitly guide reconstruction. ContextCodec adopts a dual-branch encoder that decouples acoustic details from content-focused context. The context branch is trained with a CLIP-style contrastive loss that aligns context features with phoneme indices, reducing paralinguistic leakage. During decoding, these features are injected at each decoding stage for explicit guidance. In addition, we introduce a lightweight autoregressive latent refinement module. Experiments show a strong quality-intelligibility trade-off down to 500 bps, with an RTF of 0.4886 on a typical mobile CPU.

Subject: INTERSPEECH.2026 - Speech Processing


#8 LitCodec: ASR-Guided Streaming Speech Coding with Unified Quantization [PDF] [Copy] [Kimi] [REL]

Authors: Son Dang Dinh, Nguyen Thi Minh Anh, Nhat Tran Hong, Huyen Ngo Thi Thu

Real-time speech applications demand codecs that stream with low latency while preserving both acoustic quality and linguistic content. Waveform-oriented codecs suppress fine-grained phonetic cues at low bitrates, while semantic-aware codecs rely on non-causal encoders or dual-branch architectures incompatible with streaming. We propose LitCodec, a semantic-aware neural speech codec that injects ASR-guided supervision directly into a fully causal encoder before quantization, embedding linguistic structure into a single-codebook representation via Finite Scalar Quantization. This avoids separate semantic branches, enabling single-pass streaming without sacrificing semantic fidelity. On LibriSpeech, LitCodec achieves the highest PESQ and STOI among streaming codecs (PESQ 2.56, STOI 0.925) with strong semantic preservation (WER 2.8%) at 800 bps and 50 tokens per second. At 640 bps, LitCodec maintains 3.1% WER while EnCodec degrades to 29.0% at 750 bps.

Subject: INTERSPEECH.2026 - Speech Processing


#9 An Ultra-Low-Bitrate Neural Speech Codec with Plain-to-Pseudo Synergistic Vector Quantization [PDF] [Copy] [Kimi] [REL]

Authors: Xiao-Hang Jiang, Yang Ai, Fei Liu, Rui-Chen Zheng, Jian-Qing Gao, Zhen-Hua Ling, Ji Wu

Most neural speech codecs use residual vector quantization (RVQ), in which later VQs contribute less but consume the same bitrate, leading to inefficiency. We propose P2PSynCodec, an ultra-low-bitrate neural speech codec with a plain-to-pseudo synergistic vector quantizer (P2PSVQ). P2PSVQ consists of one plain VQ and multiple pseudo VQs. The plain VQ produces basic tokens by quantization, while the pseudo VQs generate auxiliary tokens by neural prediction and incur zero transmitted bitrate. Thus, speech is decoded from the plain-VQ tokens together with predicted pseudo-VQ tokens, greatly reducing bitrate. Experiments show that P2PSynCodec achieves speech reconstruction quality comparable to competing codecs at 2.0 kbps while operating at only 0.5 kbps, demonstrating high efficiency for ultra-low-bitrate speech coding.

Subject: INTERSPEECH.2026 - Speech Processing


#10 Absorbing Discrete Diffusion for Speech Enhancement [PDF] [Copy] [Kimi] [REL]

Author: Philippe Gonzalez

Inspired by recent developments in neural speech coding and diffusion-based language modeling, we tackle speech enhancement by modeling the conditional distribution of clean speech codes given noisy speech codes using absorbing discrete diffusion. The proposed approach, which we call ADDSE, leverages both the expressive latent space of neural audio codecs and the non-autoregressive sampling procedure of diffusion models. To efficiently model the hierarchical structure of residual vector quantization codes, we propose RQDiT, which combines techniques from RQ-Transformer and diffusion Transformers for non-autoregressive modeling. Results show competitive performance in terms of non-intrusive objective metrics on two datasets, especially at low signal-to-noise ratios and with few sampling steps. Code and audio examples are available online.

Subject: INTERSPEECH.2026 - Speech Processing


#11 Schrödinger Bridge Mamba for One-Step Speech Enhancement [PDF] [Copy] [Kimi] [REL]

Authors: Jing Yang, Sirui Wang, Chao Wu, Lei Guo, Fan Fan

We present Schrödinger Bridge Mamba (SBM), a novel model for efficient speech enhancement by integrating the Schrödinger Bridge (SB) training paradigm and the Mamba architecture. Experiments of joint denoising and dereverberation tasks demonstrate SBM outperforms strong generative and discriminative methods on multiple metrics with only one step of inference while achieving a competitive real-time factor for streaming feasibility. Ablation studies reveal that the SB paradigm consistently yields improved performance across diverse architectures over conventional mapping. Furthermore, Mamba exhibits a stronger performance under the SB paradigm compared to Multi-Head Self-Attention (MHSA) and Long Short-Term Memory (LSTM) backbones. These findings highlight the synergy between the Mamba architecture and the SB trajectory-based training, providing a high-quality solution for real-world speech enhancement. Demo page: https://sbmse.github.io

Subject: INTERSPEECH.2026 - Speech Processing


#12 Speech Enhancement Based on Drifting Models [PDF] [Copy] [Kimi] [REL]

Authors: Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson

We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem. Rather than relying on iterative sampling, DriftSE natively achieves one-step inference by evolving the pushforward distribution of a mapping function to directly match the clean speech distribution. This evolution is driven by a Drifting Field, a learned correction vector that guides samples toward the high-density regions of the clean distribution, which naturally facilitates training on unpaired data by matching distributions rather than individual paired samples. We investigate the framework under two formulations: a direct mapping from the noisy observation, and a stochastic conditional generative model from a Gaussian prior. Experiments on the VoiceBank-DEMAND benchmark demonstrate that DriftSE achieves high-fidelity enhancement in a single step, outperforming multi-step diffusion baselines and establishing a new paradigm for speech enhancement.

Subject: INTERSPEECH.2026 - Speech Processing


#13 Learnable Schrödinger Bridge and Activations for Efficient Diffusion-based Speech Enhancement [PDF] [Copy] [Kimi] [REL]

Authors: Yihui Fu, Wouter Tirry, Tim Fingscheidt

Generative speech enhancement has shown significant advancements in improving speech quality in noisy environments. However, its iterative nature suffers from high inference computational complexity. In this paper, we propose EffDiffSE+, an efficient single-iteration diffusion-based speech enhancement model based on a condition DNN, a bridge DNN, and our novel learnable Schrödinger bridge (SB). Our contributions are threefold. First, we propose a single-iteration SB with Gaussian distribution initialization for the reverse process. Second, an auxiliary network is proposed to provide learnable adaptivity to the bridge DNN initial state estimation. Third, topology improvements are presented, constituting our EffDiffSE+ to consistently improve model performance. The proposed EffDiffSE+ model excels top open-source time- and frequency-domain diffusion baseline methods in PESQ, POLQA, NISQA, UTMOS, ESTOI, LPS, SBScore, SpkSim, and subjective MOS, clearly achieving an overall top rank.

Subject: INTERSPEECH.2026 - Speech Processing


#14 Mind the Gap: Detecting Cluster Exits for Robust Local Density-Based Score Normalization in Anomalous Sound Detection [PDF] [Copy] [Kimi] [REL]

Authors: Kevin Wilkinghoff, Gordon Wichern, Jonathan Le Roux, Zheng-Hua Tan

Local density-based score normalization is an effective component of distance-based embedding methods for anomalous sound detection, particularly when data densities vary across conditions or domains. In practice, however, performance depends strongly on neighborhood size. Increasing it can degrade detection accuracy when neighborhood expansion crosses cluster boundaries, violating the locality assumption of local density estimation. This observation motivates adapting the neighborhood size based on locality preservation rather than fixing it in advance. We realize this by proposing cluster exit detection, a lightweight mechanism that identifies distance discontinuities and selects neighborhood sizes accordingly. Experiments across multiple embedding models and datasets show improved robustness to neighborhood-size selection and consistent performance gains.

Subject: INTERSPEECH.2026 - Speech Processing


#15 Lung-SRAD: Spectral-Aware Regularized Audio DASS with Dual-Axis Patch-Mix Contrastive Learning for Respiratory Sound Classification [PDF] [Copy] [Kimi] [REL]

Authors: Hemansh Shridhar, Miika Toikkanen, June-Woo Kim

Recent respiratory sound classification (RSC) studies largely rely on CLS-token driven self-attention architectures such as the Audio Spectrogram Transformer (AST). While effective at modeling global context, recent analyses suggest a low-pass filtering behavior that may reduce sensitivity to localized abnormal patterns. In this work, we investigate State Space Models (SSMs) as an alternative backbone for RSC. Using the Distilled Audio State Space model, we analyze intermediate representations through spectral response curves and observe stronger preservation of mid-to-high spatial-frequency components. Based on these observations, we introduce spectral-aware layer regularization using Gaussian convolution applied to selected layers. We further propose Dual-Axis Patch-Mix contrastive learning tailored to SSM-based audio models for robust representation learning. Experiments on the ICBHI benchmark show that our approach achieves 64.48% score, outperforming the AST baseline by 5%. Code is available at https://github.com/RSC-Toolkit/Lung-SRAD.

Subject: INTERSPEECH.2026 - Speech Processing


#16 TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios [PDF] [Copy] [Kimi] [REL]

Authors: Hong Lyu, Mingru Yang, Qianhua He, Yanxiong Li, Jinxin Huang, Zhengyu Pei

There are some datasets of varying scales for audio classification (AC) applied to different tasks. However, annotated data is limited for most scenarios, such as domestic environments. To address this challenge, we propose an Automatic Audio Annotation Pipeline -TriA Pipeline, which can efficiently convert audio from various scenarios into high-quality training data with audio event annotations. A TriA dataset was constructed with the TriA Pipeline, over 2130 hours of audio covering 431 audio classes. Furthermore, we partitioned a prior-knowledge-guided subset (TriAGK) from TriA and conduct comparative experiments on three domestic AC tasks. Comparing the result on manually annotated data only and that on manually annotated data combines TriAGK, TriAGK could achieve average relative gains of 3.97% in accuracy and 3.35% in Macro-F1, validating the effectiveness of TriAGK and the TriA Pipeline.

Subject: INTERSPEECH.2026 - Speech Processing


#17 Acoustic Landmark Detector based on Conformer and HuBERT [PDF] [Copy] [Kimi] [REL]

Authors: Mateo Cámara, José Luis Blanco, Juan Ignacio Godino-Llorente, Jeung-Yoon Choi, Stefanie Shattuck-Hufnagel

Acoustic landmarks (abrupt acoustic changes tied to speech events) offer a linguistically grounded representation for speech analysis. We study automatic landmark detection with Conformer models, evaluating 14 configurations spanning architecture, loss, label representation, feature extractor, and data conditions on 1,839 manually annotated utterances with eight landmark types. We propose Gaussian soft labels with per-class temporal spread (σ=10-20 ms), improving F1@20 ms by 7.0% absolute vs. hard labels by modeling annotation variability. Frozen HuBERT features perform best without fine-tuning (F1@20 ms=0.77). Stops and fricatives are reliable (F1>0.80), while vowels remain challenging (F1≈0.55). On our corpus, our system reaches a 13.8% Landmark Error Rate (LER). This is not directly comparable to AutoLandmark (31.3%) or SpeechMark (56.5%), evaluated on a different corpus and metric. Per-class trends show detectability increases with event abruptness, consistent with Stevens' theory.

Subject: INTERSPEECH.2026 - Speech Processing


#18 Ecologically-Constrained Task Arithmetic for Multi-Taxa Bioacoustic Classifiers Without Shared Data [PDF] [Copy] [Kimi] [REL]

Authors: Ragib Amin Nihal, Benjamin Yen, Runwu Shi, Takeshi Ashizawa, Kazuhiro Nakadai

Training data for bioacoustics is scattered across taxa, regions, and institutions. Centralizing it all is often infeasible. We show that independently fine-tuned BEATs encoders can be composed into a unified 661-species classifier via task vector arithmetic without sharing data. We find that bioacoustic task vectors are near-orthogonal (cosine 0.01-0.09). Their separation aligns closely with spectral distribution distance, a gradient consistent with the acoustic niche hypothesis. This geometry makes simple averaging optimal while sign-conflict methods reduce accuracy by one to six percentage points. Composition also creates an asymmetric gap: species-rich groups lose accuracy relative to joint training while underrepresented taxa gain, a redistribution useful for equitable biodiversity monitoring. We verify linear mode connectivity across all taxonomic pairs, demonstrate zero-shot transfer to new regions, and identify domain negation as a boundary condition where composition fails. These results enable a collaborative paradigm for bioacoustics where institutions share only task vectors to assemble multi-taxa classifiers, preserving data privacy.

Subject: INTERSPEECH.2026 - Speech Processing


#19 Beyond task performance: Decoding bioacoustic embeddings with speech features [PDF] [Copy] [Kimi] [REL]

Authors: Ines Nolasco, Jules Cauzinille, Marius Miron, Gagan Narula, Milad Alizadeh, Emmanuel Fernandez, Matthieu Geist, Ellen Gilsenan-McMahon, Olivier Pietquin, Emmanuel Chemla, Sara Keen

Pretrained audio embeddings are standard in bioacoustics, yet little is known about which acoustic features they encode, nor which are useful for a given task. This limits transparency and extension to rare species or data-scarce domains. We ask which speech-like features bioacoustic embeddings encode, framing evaluation as interpretability rather than benchmarking. Using the 88 eGeMAPS features across six taxonomic groups, we apply linear and nonlinear regression probes to quantify which acoustic properties each model captures. Results confirm a "no free lunch" pattern: no single model captures the full feature space, while a concatenated embedding performs best, suggesting complementary coverage. Loudness is best encoded (R² = 0.76) while F0 is hardest to recover (R² = 0.33). Cross-referencing recoverability with per-species feature salience (NMI) offers data-driven hypotheses for model selection in bioacoustics.

Subject: INTERSPEECH.2026 - Speech Processing


#20 UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement [PDF] [Copy] [Kimi] [REL]

Authors: Haoyin Yan, Chengwei Liu, Shaofei Xue, Xiaotao Liang, Yinghao Liu, Yuxiang Kong, Zheng Xue

Neural audio codecs have largely promoted the application of language models (LMs) for speech applications. However, the effectiveness of autoregressive LM-based models in unifying speech enhancement (SE) tasks remains underexplored. In this work, we propose UniSE, a unified decoder-only LM-based framework to handle different SE tasks including speech restoration, target speaker extraction, and speech separation. Conditioned on input speech features, it autoregressively generates target discrete tokens, facilitating compatibility between distinct learning patterns of multiple tasks. To further optimize speech quality, we introduce a progressive reinforcement learning strategy with multiple assessment criteria. Experiments on several benchmarks show that UniSE achieves competitive performance compared to discriminative and generative baselines, demonstrating the capacity of LMs in unifying SE tasks. The code and demo is available at: https://github.com/alibaba/unified-audio/tree/main/QuarkAudio-UniSE.

Subject: INTERSPEECH.2026 - Speech Processing


#21 DelayGSE: A Generative Speech Enhancement Framework with Delayed Text-Aware Conditioning [PDF] [Copy] [Kimi] [REL]

Authors: Xin Yuan, Junling Lv, Zezhou Xu, Xingjun Tan, Liangliang Li, Yanqiang Lei

Recent generative speech enhancement methods based on language and diffusion models achieve strong perceptual quality but are more susceptible than discriminative approaches to speech-like hallucinations under low SNRs and transient noise. We propose DelayGSE, a text-aware generative speech enhancement framework built on a multi-codebook language model for denoising, dereverberation, and audio super-resolution. DelayGSE conditions on noisy-speech STFT features and Whisper encoder representations, and models multiple discrete codebooks in a delayed manner to stabilize generation. A text-aware mechanism suppresses hallucinations, while an importance-aware codebook weighting strategy balances perceptual fidelity and semantic consistency. Experiments demonstrate state-of-the-art performance, with ablations showing effective hallucination suppression and a 15.8% relative word error rate reduction. Audio samples are available at https://delaygse.github.io/.

Subject: INTERSPEECH.2026 - Speech Processing


#22 StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement [PDF] [Copy] [Kimi] [REL]

Authors: Xiaobin Rong, Jun Gao, Zheng Wang, Mansur Yesilbursa, Kamil Wojcicki, Jing Lu

Achieving high perceptual quality without hallucination remains a challenge in generative speech enhancement (SE). A representative approach, PASE, is robust to hallucination but has limited perceptual quality under adverse conditions. We propose StuPASE, built upon PASE to achieve studio-level quality while retaining its low-hallucination property. First, we show that finetuning PASE with dry targets rather than targets containing simulated early reflections substantially improves dereverberation. Second, to address performance limitations under strong additive noise, we replace the GAN-based generative module in PASE with a flow-matching module, enabling studio-quality generation even under highly challenging conditions. Experiments demonstrate that StuPASE consistently produces perceptually high-quality speech while maintaining low hallucination, outperforming state-of-the-art SE methods. Audio demos are available at: https://xiaobin-rong.github.io/stupase_demo/.

Subject: INTERSPEECH.2026 - Speech Processing


#23 Improving DF-Conformer using Hydra for high-fidelity generative speech enhancement on discrete codec token [PDF] [Copy] [Kimi] [REL]

Authors: Shogo Seki, Shaoxiang Dang, Li Li

The Dilated FAVOR Conformer (DF-Conformer) is an efficient variant of the Conformer architecture designed for speech enhancement (SE). It employs fast attention through positive orthogonal random features (FAVOR+) to mitigate the quadratic complexity associated with self-attention, while utilizing dilated convolution to expand the receptive field. This combination results in impressive performance across various SE models. In this paper, we propose replacing FAVOR+ with bidirectional selective structured state-space sequence models to achieve two main objectives: (1) enhancing global sequential modeling by eliminating the approximations inherent in FAVOR+, and (2) maintaining linear complexity relative to the sequence length. Specifically, we utilize Hydra, a bidirectional extension of Mamba, framed within the structured matrix mixer framework. Experiments conducted using a generative SE model on discrete codec tokens, known as Genhancer, demonstrate that the proposed method surpasses the performance of the DF-Conformer.

Subject: INTERSPEECH.2026 - Speech Processing


#24 Towards Robust Generative Speech Enhancement Using Vector Quantisation-Based Neural Audio Codec [PDF] [Copy] [Kimi] [REL]

Authors: Haixin Zhao, Nilesh Madhu

This work investigates modelling strategies in continuous and discrete latent spaces in the vector quantisation (VQ)-based neural audio codec (NAC) speech enhancement (SE), along with the role of VQ regularisation. We propose cNAC-SE and dNAC-SE frameworks that predict continuous representations and discrete tokens in latent space, respectively. Theoretical analysis and visualisations in latent space are performed to exhibit their inherent modelling mechanisms. Experimental results show that the fully fine-tuned cNAC-SE model consistently outperforms all dNAC-SE variants across diverse test conditions and achieves leading performance among established generative approaches in DNS-MOS metrics. Comparison with the discriminative counterpart shows that VQ enhances robustness through an intrinsic effect of clean-prior-constrained regularisation, independent of discrete token processing. This highlights the transferable value of VQ regularisation to other continuous modelling methods.

Subject: INTERSPEECH.2026 - Speech Processing


#25 Post-Training Speech Enhancement Language Models with Perceptual Rewards [PDF] [Copy] [Kimi] [REL]

Authors: Frédéric Berdoȥ, Luca A. Lanzendöerfer, Antonis Asonitis, Roger Wattenhofer

Speech enhancement language models achieve strong results when trained on discrete audio tokens, but their optimization relies on token-level cross-entropy rather than the perceptual metrics used for evaluation. We introduce a post-training stage for autoregressive speech enhancement language models using Group Sequence Policy Optimization (GSPO) with multi-metric perceptual rewards. Our method directly optimizes non-differentiable quality metrics (DNSMOS, WER, and UTMOS) as reward signals, without learned surrogates or offline preference pairs. Applied to two autoregressive base models, UniSE and GenSE, our approach achieves state-of-the-art results on the DNS2020 benchmark. A human evaluation ablation further shows that the composite multi-metric reward is preferred over any single-metric variant, confirming that multi-reward optimization avoids the reward hacking observed with single-metric training.

Subject: INTERSPEECH.2026 - Speech Processing