yen26@interspeech_2026@ISCA

Total: 1

#1 MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding [PDF] [Copy] [Kimi] [REL]

Authors: Hao Yen, Pin-Jui Ku, Ante Jukić, Sabato Marco Siniscalchi

In sequence-to-sequence Transformer-based automatic speech recognition (ASR), autoregressive (AR) models achieve strong accuracy but suffer from slow decoding, while non-autoregressive (NAR) models enable parallel decoding at the cost of degraded performance. We propose a principled NAR ASR framework based on Masked Diffusion Models to significantly reduce this gap. A pre-trained speech encoder is coupled with a Transformer diffusion decoder conditioned on acoustic features and partially masked transcripts for parallel token prediction. To mitigate the training-inference mismatch, we introduce Iterative Self-Correction Training that exposes the model to its own intermediate predictions. We also propose a Position-Biased Entropy-Bounded Confidence-based sampler to further boost results. Experiments across multiple benchmarks demonstrate consistent gains over prior NAR models and competitive performance with strong AR baselines, while retaining parallel decoding efficiency.

Subject: INTERSPEECH.2026 - Speech Recognition