Real-Time Voice Activity Detection using Hybrid Deep Learning Models with a Multi-Noise Detector
Aliouat Mahfoud, Mohamed Djendi · 2025
In this study, we present a novel voice activity detection (VAD) system leveraging deep learning techniques to enhance speech detection in noisy environments. Traditional approaches, such as energy-based methods and SNR estimation, have been widely employed to address challenges in speech applications. However, recent advancements in artificial intelligence and the availability of extensive datasets have enabled more robust solutions for speech enhancement and noise reduction. The proposed system incorporates a two-stage architecture. In the first stage, a Bi-LSTM-based noise detection model is utilized to identify the type of noise present. The second stage employs a set of noise-specific VAD models built on hybrid BiLSTM-GRU architecture, tailored to maintain high performance across various acoustic environments. The system was evaluated using multiple performance metrics, including accuracy, F1 score, and recall, under diverse noise types and signal-to-noise ratio (SNR) levels. Results demonstrate that the Bi-LSTM-GRU model outperforms existing methods, achieving superior accuracy across all noise conditions. Notably, for Volvo noise, the model achieved an F1 score of 97.92% and a recall of 98.44%, indicating its ability to effectively handle complex noisy environments. These findings highlight the resilience of the proposed system in accurately detecting voice activity across a wide range of noise types and intensities, offering promising potential for real-world applications in communication systems, industrial settings, and beyond.