Robust Voice Activity Detection Based on Complementary BLSTM Enhancement Stage
Iman Shahryary, Sanaz Seyedin, Seyed Mohammad Ahadi · 2020
In this paper, we propose a new two-stage deep structure with a joint learning technique to improve Voice Activity Detection (VAD) in different noisy conditions especially in unseen noises. The first stage of our proposed method deals with the enhancement of the noisy signal, which is complementary to the second stage. Bidirectional Long Short-Term Memory (BLSTM) architecture is used in this part so as to take benefit from both previous and upcoming frames. The second stage uses the enhanced frames features to predict the speech presence probability. Based on previous studies, we use Multi-Resolution Cochleagram (MRCG) features to achieve higher robustness. We evaluate our proposed method using the Area Under the Curve (AUC) and precision metrics in TIMIT corpus. Based on our evaluations, the proposed method outperforms other state-of-the-art methods based on deep structures as baseline, both in AUC and precision metrics. The proposed method's AUC improvement versus other methods, in noises not seen in the training step, is significant.