Improving Voice Activity Detection by using Denoising-Based Techniques with Convolutional LSTM
Nattapong Kurpukdee, Surasak Boonkla, Vataya Chunwijitras, Phuttapong Sertsi, Sawit Kasuriya · 2019
The performance of voice activity detection (VAD) is drastically degraded when observed speech signals are from unseen noisy environments. In this paper, we propose denoisingbased VAD to cope with the unseen noises. The proposed VAD system mainly consists of two stages for denoising and speech/non-speech classification. In the first stage, either logmagnitude spectral estimator (LSA) or convolutional long shortterm memory neural network autoencoder (CLAE) is applied to eliminate the noises. The convolutional bidirectional long-shortterm memory deep neural network (CBLDNN) is employed for the speech/non-speech classification. The results showed that the proposed VAD was better than the baseline. Furthermore, our CLAE tends to outperform the LSA in denoising algorithms when the signal-to-noise ratio is 5dB.