Speech Activity Detection in Naturalistic Audio Environments: Fearless Steps Apollo Corpus

Lakshmish Kaushik, Abhijeet Sangwan, John H. L. Hansen · IEEE Signal Processing Letters · 2018

Speech activity detection (SAD) is a fundamental building block for most spoken language technology systems. Developing efficient SAD systems in highly naturalist data scenarios is a challenge. In this study, we investigate the SAD problem on NASAs Apollo space mission data [1]. Apollo data consists of long-term naturalistic audio recordings (i.e., 6-12 day missions). The Apollo data poses various challenges like: 1) noise distortion with variable SNR, 2) channel distortion, 3) very high density of speech, 4) foreground versus background speech, and 5) extended periods of nonspeech activity. In this study, we use threshold optimized combo-SAD [21] as our baseline unsupervised system. This technique was developed to address variable speech/nonspeech density issues in long-term audio data. To mitigate issues related to Apollo audio loops, multispeaker scenarios including foreground versus background conversations within loops, and highly noisy background, a new curriculum learning (CL) based convolutional neural network (CNN) model is developed. This efficient method leverages the long-term learning capability of CNN and CL strategies where data are trained in a manner that improves the efficiency during the learning process. Here, we use signal-to-noise ratio as the learning parameter. Our experiments on free flowing Apollo audio data show that the proposed approach provides a significant improvement in SAD performance (> 10%).

Read the paper · More papers on PaperTik