Combining Deep Neural Networks and Beamforming for Real-Time Multi-Channel Speech Enhancement using a Wireless Acoustic Sensor Network

Enea Ceolini, Shih‐Chii Liu · 2019

This work presents a multi-channel speech enhancement algorithm using a neural network combined with beamforming deployed realtime on a wireless acoustic sensor network (WASN) of distributed microphones. We combine spectral mask estimation via a deep neural network together with spatial filtering to obtain a robust speech enhancement system even in difficult real-world scenarios (e.g. speech in noise, reverberant environments). Although the model is trained on simulated data, it performs comparably well on real-world tasks relative to an ideal oracle beamformer. We show that the model can be deployed on a WASN platform that allows for remote placement of microphones and on-board computing. We consider models with a small parameter count and low computational complexity. It achieves signal-to-distortion ratio (SDR) improvements of up to 10dB in a real-world scenario and runs real-time on-board the WASN, with a latency in the order of hundreds of milliseconds.

Read the paper · More papers on PaperTik