Human Speech Extraction from Composite Audio Signal in Real-Time using Deep Neural Network

Pavan Nettam, R Rakshitha, Meghana Nekar, Amitram Achunala, S. S. Shylaja, Preethi Anantharaman · 2024

Studies indicate that humans have the exceptional capability to focus on speech even in environments with complex background noises. To create a computational equivalent of this ability, recent studies have focused on developing speech extraction algorithms. These algorithms aim to extract human speech signals from a mixture of overlapping background sounds. In this paper, efforts are made to utilize the novel encoder-decoder neural network architecture to accomplish this task for real-time and streaming applications. The training of the model involved utilizing synthetic 6-second single-channel composite audio mixtures generated from three distinct datasets using the Scaper toolkit. Our results showcase an impressive SI-SNRi of 14.209 and a latency of 874 ms, obtained on a consumer-grade Intel Quad-Core i5 processor with default multithreading. The obtained outcomes provide compelling evidence for the feasibility and efficacy of the proposed approach, establishing a strong proof-of-concept for the real-world application of the model. Streaming human speech extraction could revolutionize acoustic applications for headphones, hearing aids, and telephony.

Read the paper · More papers on PaperTik