Speech/nonspeech discrimination in reverberant teleconferencing environments for acoustical source location

Christopher V. Alvino, Deborah M. Grove, J.T. Kleban, Shashank Sathyanarayana · The Journal of the Acoustical Society of America · 2000

A real-time acoustical source location system implemented at the CAIP Center at Rutgers University presently tracks all sounds within a room. The system aims a video camera and audio sensors at the determined sound source for transmission of video and audio information to a remote teleconferencing location. It is desired to only direct attention to speech and to ignore nonspeech such as the shuffling of feet or the moving of papers. Speech detection is implemented to allow the system to aim sensors only at talkers. Present detection methods depend upon energy levels and the pitch present in voiced sounds to detect speech. In a reverberant environment the accuracy of speech detection degrades. Speech detection accuracy in reverberation is tested. Measures of false acceptance of nonspeech signals and false rejection of speech signals are used to determine which speech detection methods are optimal in different reverberant rooms. Speech and nonspeech in reverberant rooms are simulated using impulse responses generated by CATT-Acoustic. The test environments range from anechoic to heavily reverberant. A look-back buffer and adaptive tracking of energy and zero-crossing levels are used to allow speech detection to work optimally with the current source location system.

Read the paper · More papers on PaperTik