A Method for Voice Activity Detection using K-Means Clustering
Atul Rohit Agarwal, Sourabh Tiwari, Vinay Vasanth Patage, S. Sankar Ganesh, M. Sudhakar · 2022 13th International Conference on Computing Communication and Networking Technologies (ICCCNT) · 2022
Human-Machine interaction through voice modality in recent time has resulted in both research and business use cases. Major business organizations are making the shift towards developing their state-of-the-art voice assistants, whose accuracy of understanding human voice command, is not affected to surrounding noise. One of the first steps towards developing such an interactive model is understanding and segmenting voice activity regions in an audio clip from all other noise present. In this paper we propose the use of a weighted clustering approach to solve this problem. The proposed solution utilizes four audio features, Mel frequency cepstral coefficients (MFCC), Spectral Roll-Off, Spectral Centroid and Zero Crossing Rate for every 0.125 second sub segment of the audio clip. Next, these sub segments are clustered using K Means Clustering algorithm into 2 clusters: Voice Activity and Noise Activity. This approach and solution, provides a simple and lightweight solution to the voice activity detection problem.