Joint Voice Activity Detection and Quality Estimation for Efficient Speech Preprocessing
Nikita Khmelev, Alexandr Anikin, Anastasia Zorkina, Anastasia Korenevskaya, Sergey Arkad'yevich Novoselov, Stepan Malykh, Vladimir Volokhov, Vladislav Marchevskiy, Marina Volkova, Galina Lavrentyeva · 2025
The efficiency of automatic speaker verification (ASV) and diarization systems is highly dependent on the quality of input speech pre-processing, especially under adverse acoustic environments. This paper introduces a novel neural network-based approach that integrates voice activity detection (VAD) and speech quality estimation (QE) to enhance system robustness across various testing conditions. We evaluate the proposed method against existing VAD sys-tems (Brouhaha, Silero, SpeechBrain, WebRTC), analyzing their performance in speech detection tasks. Furthermore, we assess the effectiveness of integrated QE algorithms in estimating key signal quality parameters, including reverberation time (RT60) and signal-to-noise ratio (SNR). Experimental results confirm that incorporating the proposed VAD and QE into ASV and diarization systems improves overall performance across various acoustic scenarios, underscoring their significance for real-world speech processing applications.