Active source location and beamforming
Ramani Duraiswami, Dmitry N. Zotkin, Eugene Borovikov, Larry Steven Davis · The Journal of the Acoustical Society of America · 2000
A fundamental problem in designing user interfaces that incorporate speech and vision is locating the user(s) and enhancing the quality of their speech. We present a novel algorithm for actively doing this. Two complementary coarse-to-fine strategies that divide both the space being searched and the received signal in the frequency domain are used to arrive at an efficient dynamic search algorithm that is capable of both locating and beamforming multiple sources distributed in space. The algorithm makes no assumptions on noise statistics, and has a deterministic performance bound. It is able to make use of prior information on the possible location of sources obtained via means such as video. The algorithm is particularly applicable to locating speech sources in rooms of regular size, and is thus suitable for videoconferencing and user-interface applications. The algorithm and details are presented. [Support from the W. M. Keck Foundation and ONR Contract No. N000149510521 is gratefully acknowledged.]