Total: 1
For open-vocabulary keyword spotting, phoneme-level alignment has been introduced to improve the performance on acoustically confusable words. However, most existing studies utilize non-streaming methods, which are unsuitable for streaming scenarios. Recent methods achieve streaming phoneme alignment via connectionist temporal classification (CTC) algorithms, but they are limited to text-only enrollment. To address these issues, we propose a novel streaming multimodal open-vocabulary keyword spotting method based on phoneme-level alignment. This method achieves fine-grained modeling with training-inference consistency through the W-CTC forced alignment algorithm and multimodal phoneme-level contrastive learning. Additionally, we propose a data augmentation method based on CTC beam-search to dynamically mine hard negative samples. Experiments on the LibriPhrase dataset demonstrate that the proposed method achieves the best results.