Spike-VAD: Efficient and Robust Spiking Neural Network for Voice Activity Detection
Kexin Shi, Mengshu Hou, Xiaoling Luo, Dehao Zhang, Hanwen Liu, Jingya Wang · IEEE Transactions on Cognitive and Developmental Systems · 2025
In modern speech applications, achieving both low power consumption and noise robustness is critically essential. A well-designed Voice Activity Detection (VAD) front-end minimizes processing demands. Spiking Neural Networks (SNNs), a cutting-edge paradigm in brain-inspired computing, excel in energy efficiency due to their spike-based processing mechanisms. This makes them a promising solution for developing more efficient VAD models. In recent years, researchers have achieved notable advancements in applying SNNs to VAD, particularly in energy efficiency and performance. However, current SNN-based VAD models still struggle to achieve sufficient robustness and fail to fully exploit the low-power potential of SNNs. To address this challenge, we propose an energy-efficient and highly robust spike-based VAD model, called Spike-VAD. Spike-VAD leverages an energy-saving resonate-and-fire frequency (RF-FRE) spike encoding scheme, eliminating the need for computation-intensive Fourier Transform (FT) operations. Inspired by the human auditory frequency preference, a spike-based attention module is designed to refine the encoded spike features and enhance robustness. Furthermore, an Adaptive Memory Modulation Strategy (AMMS) is introduced to dynamically modulate historical information from the audio, facilitating more effective decision-making. Experiments on the QUT-NOISE-TIMIT dataset indicate that compared with previous SNN-based VAD models, our model achieves state-of-the-art (SOTA) performance in both robustness and energy consumption.