choi26@interspeech_2026@ISCA

Total: 1

#1 Systematic PTQ Study of Integer and Floating-Point Formats for On-Device Whisper ASR [PDF] [Copy] [Kimi] [REL]

Authors: Woosuk Choi, Dohyeon Lee, Taehyung Kim, Hyukjun Lee

Whisper and similar encoder-decoder ASR models are increasingly deployed on mobile and edge devices, yet it remains unclear how quantization format choices affect their accuracy in practice. We systematically evaluate post-training quantization (PTQ) for Whisper (tiny.en,base.en) across 80+ configurations covering INT8/4/3 and FP8/FP4/NVFP4/MXFP4. Our main finding is that activation bit-width matters far more than weight format: dropping activations from 16-bit to 8-bit costs 1-3% absolute WER, while INT16 and FP16 activations are indistinguishable in accuracy. Among 4-bit formats, NVFP4 W4A16 comes within 0.07% of full-precision at 6.4× compression, whereas MXFP4 fails under standard PTQ and needs additional fixes. FP activation paths are also the better hardware choice, since FP multipliers are more area-efficient than INT at the same bit-width. We close with a Pareto analysis and six practical guidelines for format selection across memory budgets from 20 to 80 MB.

Subject: INTERSPEECH.2026 - Speech Recognition