Speaker Position Detection System Using Audio-visual Information
VNaoshi Matsuo · 1999
This paper describes a speaker position detection system that achieves a high degree of accuracy using a multimodal interface that integrates audio and visual information from a microphone array and a camera. First, the system detects the position and angle of the microphone array relative to the camera position. Next, audio processing detects sound source positions and visual processing detects the positions of human faces. Finally, by integrating the sound source positions and face positions, the system determines the speaker’s position. The system can integrate audio and visual information, even if the spatial relationship between the microphone array and camera is initially unknown. The system achieves a high detection rate for the speaker’s position in a noisy environment.