Embedded Discriminant Analysis based Speech Activity Detection for Unsupervised Stress Speech Clustering

Barlian Henryranu Prasetio, Hiroki Tamura, Koichi Tanno · 2020

Speech activity detection (SAD) or sometimes called voice activity detection (VAD), is a crucial part of most speech-related applications. The SAD system serves to ensure the primary system processes only speech segments. Many speech-based systems have reported that detection accuracy is thanks to the robustness of their SAD system. Various SAD methods have been explored and enhanced in addressing noisy environments, but a few of them notice the emotional condition of the speakers. Whereas, in real conditions, emotions (such as stress) can pose a considerable impact on SAD system performance. In this paper, we propose a compact SAD system that is able to harmonize with the altered speech characteristics due to the presence of emotion and also powerful in high noise conditions. Since there is a similarity between emotional effect and channel effect, the advantages of the proposed SAD system is the applied of a new channel compensation scheme (termed as embedded discriminant analysis, EDA) that works in the i-vector space. We design the EDA in such a way so that it could compensate the presence of emotional condition. EDA transforms original i-vector to a lower-dimensional denoise embedding space. We develop EDA as simple and efficient as the linear discriminant analysis (LDA). The cosine similarity algorithm is applied to calculate the resemblance score between the audio target and the speech/non-speech models, and also for deciding the decision threshold. The effectiveness of the proposed SAD system is evaluated in the clustering task of Speech Under Simulated and Actual Stress (SUSAS) data, that aimed for the stress speech clustering (SSC) system. Contribution-We propose a SAD system which not only strong in noisy environments but also be able to compensate the presence of emotional conditions.

Read the paper · More papers on PaperTik