Total: 1
In recent years, attention-based multi-modal open-vocabulary keyword spotting (KWS) has attracted attention, yet its streaming deployment on device with such a design remains underexplored. We identify a data flow mismatch in traditional frameworks: speech as Key/Value requires global context but, in streaming, provides only local frames; enrollment representation as Query supplies only local information, despite its inherent global semantics. To resolve this mismatch, we propose a role-swapping streaming open-vocabulary KWS approach: streaming speech becomes Query (carrying frame-level local info) and enrollment representation as Key/Value (encoding global semantics), enabling streaming processing. Without loss of generality, we instantiate this framework with a text-registered model (~0.8M parameters) and evaluate it on LibriPhrase, achieving an EER/AUC of 6.82%/97.95% on easy negative subset and 28.21%/79.19% on hard negative subset.