EEMD and Double Thresholds Integrated Voice Activity Detection
Shan Meng, Mijit Ablimit, Askar Hamdulla · 2022
Voice activity detection (VAD) is an important preprocessing for voice applications. Anti-noise performance is the most important evaluation index of VAD algorithm. The traditional dual-threshold-based VAD algorithm has very low detection accuracy in a low signal-to-noise ratio environment. This paper proposes a voice activity detection algorithm based on Ensemble Empirical Mode Decomposition (EEMD) combined with the dual-threshold method, which integrates the decomposition of EEMD. The denoising feature is combined with VAD based on dual thresholds, and dual thresholds are set for VAD to improve the anti-noise performance and accuracy of the algorithm. VAD is divided into three categories: based on feature parameters, based on pattern recognition, and based on deep learning. The VAD algorithms based on pattern recognition and deep learning require the support of big training data to achieve good detection results. And it is too complex and requires a large amount of computation, thus its application and real-time deployment have been greatly affected. The traditional VAD methods based on feature parameters only needs to use the short-term energy and short-term zero-crossing rate as the judgment criteria for activity detection. Thise algorithm is simple, easy to deploy, and applicable even if it is dozens of voices. At the same time, the EEMD changes the extreme point characteristics of the signal by adding different white noises of the same amplitude each time, and then performs the overall average of the corresponding IMF(intrinsic mode function) obtained by multiple EMDs to offset the added white noise, therefor effectively suppressing the mode-mixing. Thus the production of state aliasing can better improve the anti-noise performance of the VAD algorithm.