zhang26fa@interspeech_2026@ISCA

Total: 1

#1 MPA-KWS: Multi-Modal Phoneme-Level Alignment for Streaming Open-Vocabulary Keyword Spotting [PDF] [Copy] [Kimi] [REL]

Authors: Jue Zhang, Guibin Zheng, Jiarui Zhang, Jiqing Han, Chenhao Jing

For open-vocabulary keyword spotting, phoneme-level alignment has been introduced to improve the performance on acoustically confusable words. However, most existing studies utilize non-streaming methods, which are unsuitable for streaming scenarios. Recent methods achieve streaming phoneme alignment via connectionist temporal classification (CTC) algorithms, but they are limited to text-only enrollment. To address these issues, we propose a novel streaming multimodal open-vocabulary keyword spotting method based on phoneme-level alignment. This method achieves fine-grained modeling with training-inference consistency through the W-CTC forced alignment algorithm and multimodal phoneme-level contrastive learning. Additionally, we propose a data augmentation method based on CTC beam-search to dynamically mine hard negative samples. Experiments on the LibriPhrase dataset demonstrate that the proposed method achieves the best results.

Subject: INTERSPEECH.2026 - Others