Total: 1
Post-training quantization (PTQ) methods such as GPTQ and AWQ compress large language models effectively, but applying them to automatic speech recognition (ASR) at ultra-low bit-widths (2-3 bits) causes severe transcription degradation. Unlike discrete text tokens, speech activations are highly correlated in steady-state and zero-padded regions yet change rapidly at phonetic boundaries; when all frames contribute equally to the Hessian, static frames dominate calibration and obscure informative transitions. We propose DiffAQ, which computes frame-to-frame activation differences to measure the rate of acoustic change and assigns Hessian importance proportionally, concentrating quantization precision on phonetically critical frames. As a training-free modification to GPTQ, DiffAQ consistently reduces WER across various Whisper sizes and standard benchmarks, with the largest gains at 2-bit precision where baseline methods frequently produce degenerate outputs.