Voice Activity Detection based on Multi-Dilated Convolutional Neural Network
Jaeseok Kim, Heejin Choi, Jinuk Park, Juntae Kim, Minsoo Hahn · 2018
Voice activity detection (VAD) is important frontend of audio signal processing. Therefore, VAD system is required to have high performance with low computation cost. Contextual information (CI). CI is useful for improving performance in low signal-to-noise scenarios. We presents to apply Convolutional neural networks (CNN) for exploiting CI. Because CNN can efficiently control the duration of CI with convolution filter. Conventional CNN needs need a large number of parameters and computation cost when employing long-term CI. We apply the dilated convolution to utilize the long-term CI using few numbers of parameters. Dilated convolution is a method configuring subsampled large filter. Since the speech signal has a high correlation between adjacent frames, it is effective to use subsampled input signal with applying dilated convolution. We propose multi-dilated convolution to employing multiple CI. The multi-Dilated convolution layer extracts features of multiple CI by consisting filters having different filter size. We used TIMIT dataset for the speech data. We used 'A sound effect library' and Noise-X for noise data. Experimental result show the VAD system applied multi-dilated convolution layer has higher average performance than reference systems