jeon26c@interspeech_2026@ISCA

Total: 1

#1 Not All Frames Are Equal: Difference-Aware Quantization for Ultra-Low-Bit ASR [PDF] [Copy] [Kimi] [REL]

Authors: Woori Jeon, Jungmin So

Post-training quantization (PTQ) methods such as GPTQ and AWQ compress large language models effectively, but applying them to automatic speech recognition (ASR) at ultra-low bit-widths (2-3 bits) causes severe transcription degradation. Unlike discrete text tokens, speech activations are highly correlated in steady-state and zero-padded regions yet change rapidly at phonetic boundaries; when all frames contribute equally to the Hessian, static frames dominate calibration and obscure informative transitions. We propose DiffAQ, which computes frame-to-frame activation differences to measure the rate of acoustic change and assigns Hessian importance proportionally, concentrating quantization precision on phonetically critical frames. As a training-free modification to GPTQ, DiffAQ consistently reduces WER across various Whisper sizes and standard benchmarks, with the largest gains at 2-bit precision where baseline methods frequently produce degenerate outputs.

Subject: INTERSPEECH.2026 - Speech Recognition