baek26@interspeech_2026@ISCA

Total: 1

#1 SPARK: Efficient Audio-Text Matching for User-Defined Keyword Spotting via Spiking Neural Networks [PDF] [Copy] [Kimi] [REL]

Authors: Seung-Yeop Baek, Sangho Han, Joon-Hyuk Chang

With the increasing prevalence of voice-driven interaction, keyword spotting (KWS) has become an essential component of hands-free control. This has led to a growing demand for user-defined KWS, allowing users to customize target keywords via text. While various models have emerged to support this, their high computational costs and energy consumption make real-world deployment challenging. In response, we introduce SPARK, a SPike-driven Audio-text matching framework for eneRgy-efficient user-defined Keyword spotting. SPARK employs a spike-driven attention mechanism to enable end-to-end processing within the spiking domain, replacing heavy floating-point operations with low-cost accumulate operations. Our experimental results on the LibriPhrase test dataset demonstrate that SPARK achieves competitive performance with a 2.1 times reduction in parameter count and a 21.7 times reduction in energy consumption compared to its artificial neural network counterpart.

Subject: INTERSPEECH.2026 - Others