wu26k@interspeech_2026@ISCA

Total: 1

#1 Leveraging Temporal Redundancy via Layer-wise Key-Value Pooling Attention for Efficient ASR [PDF] [Copy] [Kimi] [REL]

Authors: Yi Wu, Guibin Zheng, Chenhao Jing, Jiqing Han, Jiarui Zhang

Transformer-based models have revolutionized Automatic Speech Recognition. However, the quadratic complexity of self-attention and feature redundancy necessitate deep stacking, limiting long-sequence efficiency. To address this, we propose Key-Value Pooling Attention (KV-Pooling), which leverages high temporal redundancy in speech by average pooling to compress Key/Value tensors, smoothing feature distribution spikes while retaining Query resolution to maintain modeling precision. This reduces reliance on network depth. Building on this, we introduce KV-Pooling-Zipformer, integrating this mechanism into Zipformer by pre-setting differential pooling strides based on layer-wise feature abstraction. Compared to the native Zipformer, RNN-T experiments show absolute reductions of 0.2% CER on AISHELL-1 and 0.3% WER on LibriSpeech. Notably, it improves inference RTF by 10% on an AMD EPYC 7763, achieving synergistic gains in accuracy and efficiency.

Subject: INTERSPEECH.2026 - Speech Recognition