The robust feature extraction of audio signal by using VGGish model

Mandar Pramod Diwakar, Brijendra Parasnath Gupta · Research Square · 2023

Abstract This research paper explores the use of the VGGish pre-trained model for feature extraction in the context of speech enhancement. The objective is to investigate the effectiveness of VGGish in capturing relevant speech features that can be utilized to enhance speech quality and reduce noise interference. The experimentation is conducted on the MUSAN dataset, and the results demonstrate the capability of the VGGish model in extracting rich and discriminative features encompassing spectral, temporal, and perceptual characteristics of speech. These features are then employed in various speech enhancement techniques to improve speech intelligibility, enhance spectral clarity, and reduce artifacts caused by noise and distortions. Comparative analysis with traditional methods reveals the superior performance of the VGGish model in capturing a comprehensive representation of the speech signal, leading to better discrimination between speech and noise components. The findings highlight the potential of the VGGish model for speech enhancement applications, offering opportunities for improved communication systems, automatic speech recognition, and audio processing in diverse domains. Future research directions include optimizing the VGGish model for specific speech enhancement tasks, exploring novel feature fusion techniques, and integrating other deep learning architectures to further enhance system performance and flexibility. Overall, this research contributes to advancing speech processing and provides a foundation for enhancing speech quality, reducing noise interference, and improving the overall listening experience.

Read the paper · More papers on PaperTik