A Hybrid CRNN for Robust Voice Activity Detection in Noisy Environments
Rim Boukri, Atef Farrouki · 2025
Voice Activity Detection (VAD) is an important tool used as a preliminary task in most speech applications. The main goal of VAD is to detect speech segments and then discard unwanted silent intervals to minimize bandwidth and energy consumption in discontinuous transmission model. In this paper, we propose a novel hybrid model for VAD utilizing the Modified Discrete Cosine Transform (MDCT) power spectrum as input features and a Convolutional Recurrent Neural Network (CRNN) model under non-stationary and highly noise situations. This hybrid CRNN model combines a Dilated convolutional neural network (CNN), a bidirectional Long-Short-Term Memory (bi_LSTM) and a Gated Recurrent Unit (GRU) into one unified architecture. This combination leverages the strengths of each component: CNNs effectively extract local features, while bi_LSTM and GRU networks enhance temporal modeling, making their association more suitable for VAD than each individual network independently used. The performances of the proposed model have been evaluated via TIMIT and ST-CMDS Chinese datasets in stationary and non-stationary noisy environments. The proposed technique achieves an accuracy of 99.74% on both clean and low stationary noisy conditions and 95.05% in the presence of transients. The results show that the hybrid CRNN model for VAD acts robustly against non-stationary and strongly noisy situations, comparatively to recent proposed methods.