Total: 1
Cross-attention is widely used in speech-to-text (S2T) systems and often exploited for downstream applications such as timestamp prediction and speech-text alignment, under the assumption that it reflects input-output dependencies. While extensively debated in NLP, its explanatory role remains underexplored in the speech domain. We empirically assess the explanatory power of cross-attention in S2T models by comparing attention scores with input saliency maps from feature-attribution methods. Our analysis spans monolingual and multilingual, single-task and multi-task models at multiple scales. We find moderate alignment between attention and saliency, particularly when aggregating across heads and layers. However, cross-attention captures only about 50% of input relevance and, at best, 52-75% of the encoder saliency. These results show that cross-attention offers useful but incomplete explanatory cues and should be interpreted with caution as a proxy for model behavior in S2T systems.