Total: 1
Transformer-based models have revolutionized Automatic Speech Recognition. However, the quadratic complexity of self-attention and feature redundancy necessitate deep stacking, limiting long-sequence efficiency. To address this, we propose Key-Value Pooling Attention (KV-Pooling), which leverages high temporal redundancy in speech by average pooling to compress Key/Value tensors, smoothing feature distribution spikes while retaining Query resolution to maintain modeling precision. This reduces reliance on network depth. Building on this, we introduce KV-Pooling-Zipformer, integrating this mechanism into Zipformer by pre-setting differential pooling strides based on layer-wise feature abstraction. Compared to the native Zipformer, RNN-T experiments show absolute reductions of 0.2% CER on AISHELL-1 and 0.3% WER on LibriSpeech. Notably, it improves inference RTF by 10% on an AMD EPYC 7763, achieving synergistic gains in accuracy and efficiency.