Eyes and ears: Attentive teleconferencing utilizing audio and video cues
Bill Kapralos, Michael R. M. Jenkin, John K. Tsotsos, Evangelos Milios · The Journal of the Acoustical Society of America · 2001
The multiple speaker teleconferencing systems currently available typically focus on a single speaker and provide limited, if any, automatic speaker tracking technologies. However, in a multiple-speaker setting, speakers must be localized and tracked in both the video and audio domains. Although many fast and portable video trackers capable of locating and tracking humans exist, they employ conventional cameras thereby providing a narrow field of view. In addition, audio localization systems are expensive, nonportable, and computationally intensive. Furthermore, there have been very few attempts to combine both audio and visual systems. This work investigates the development of a simple, economical, and compact teleconferencing system utilizing both audio and video cues. An omni-directional video sensor is used to provide a view of the entire visual hemisphere thereby providing multiple dynamic views of all participants. Using a statistical color model and simple geometrical properties, the location of each participant’s face is determined and provided to the audio system as a possible direction to a sound source. Beam forming with a small, compact microphone array allows the audio system to detect and focus on the speech of each participant. The results of experiments conducted in normal, reverberant environments indicate the effectiveness of the system.