A binaural model using equalization/cancellation and simulated head movements to localize and extract one speaker from a mixture

Nikhil Deshpande, Jonas Braasch · The Journal of the Acoustical Society of America · 2016

This model takes a mixture of two simultaneous speech signals at unique azimuth positions to extract either speaker using the equalization/cancelation (EC) method. Head-related transfer functions are used to spatialize the two sound sources. The model localizes the sources by analyzing interaural time differences and then virtually rotates its head to find the position for the best signal-to-noise ratio. Next, the model segments the mixed speech signal in time and frequency bins, and uses an EC algorithm in each bin to compensate the target signal from the mixture. From the residual non-cancelled energy, it generates a binary map and overlays this on the spectrogram. The ability of the model to cancel out the target signal determines the bins where the target is actually present. The signal is then reconstructed in time and frequency, leaving only one desired target signal. The model achieves signal-to-noise ratios of up to 80 dB. [This material is based upon work supported by the National Science Foundation under Grant Nos. 1320059 and 1539276.]

Read the paper · More papers on PaperTik