Total: 1
With the increasing prevalence of voice-driven interaction, keyword spotting (KWS) has become an essential component of hands-free control. This has led to a growing demand for user-defined KWS, allowing users to customize target keywords via text. While various models have emerged to support this, their high computational costs and energy consumption make real-world deployment challenging. In response, we introduce SPARK, a SPike-driven Audio-text matching framework for eneRgy-efficient user-defined Keyword spotting. SPARK employs a spike-driven attention mechanism to enable end-to-end processing within the spiking domain, replacing heavy floating-point operations with low-cost accumulate operations. Our experimental results on the LibriPhrase test dataset demonstrate that SPARK achieves competitive performance with a 2.1 times reduction in parameter count and a 21.7 times reduction in energy consumption compared to its artificial neural network counterpart.