Speech localization in any direction using power and frequency signatures, gradients, and differences.

Colin L. Barnhill, James E. West · The Journal of the Acoustical Society of America · 2009

Speech localization is a relatively simple task for a human but, often, a difficult task for acoustic arrays. Although arrays can localize impulsive and narrow-band sources through the use of cross-correlation and subspace methods (like MUSIC), these methods break down when there is low SNR, reverberant conditions, or the source is neither impulsive nor narrow-band. Interesting speech sources (for teleconferencing or surveillance) are most often in environments where many of the breakdown conditions exist. New or multiple methods are necessary to reliably localize speech. A new algorithm will be presented that makes use of multiple localization methods to locate the desired speech signal. The localization methods are based on the time and frequency power signatures, gradients, and differences of third order supercardioid beams. Multiple beams are used to spatially segregate power signatures, and the results are clustered to determine position. Using a spherical array, it is possible to localize speech in any direction using these beams. Experimental acoustic and mathematical results in real room situations will be presented.

Read the paper · More papers on PaperTik