A Fusion Model for Robust Voice Activity Detection

Guan-Bo Wang, Wei-Qiang Zhang · 2019

As an indispensable front-end system, it is crucial for voice activity detection (VAD) system to be robust in all kinds of conditions. In this paper, we propose a fusion model based VAD system. A supervised fusing strategy is introduced to improve system performance in diverse data domains. We evaluate our proposed system on development datasets of Public Safety Communications (PSC), Video Annotation for Speech Technologies (VAST) and Babel from NIST Open Speech Analytic Technologies 2019 (OpenSAT19), each of which has its own challenges for VAD systems. Experimental results show the robustness of our fusion model. Compared to the baseline system, our proposed system achieves better performance under the OpenSAT19 official evaluation metrics in all three datasets.

Read the paper · More papers on PaperTik