Source-blind localization and segregation model featuring head movement and reflection removal

Nikhil Deshpande, Jonas Braasch · The Journal of the Acoustical Society of America · 2017

This model takes two simultaneous speech signals, spatialized to unique azimuth positions and convolved with a simple multi-tap stereo impulse response. The model first identifies reflections and generates an inversion filter for the left and right channels. It then localizes the sources and virtually rotates its head to a known orientation for the best resulting segregation of the sources. Next, the model segments the input signals in time and frequency, applies the inverse filter, and searches for residual energy in each bin to compensate the target signal from the mixture. From the residual non-canceled energy, it generates a binary masking map and overlays this on the mixed signal’s spectrogram to extract only the target signal. Improvement in SNR from head rotation approaches over 30 dB.

Read the paper · More papers on PaperTik