GENERALIZED POSTERIOR PROBABILITY FOR VERIFYING RECOGNIZED WORDS OPTIMALLY IN MICROPHONE ARRAY APPLICATIONS
Wolfgang Herbordt, Frank K. Soong, Satoshi Nakamura · 2005
In a large vocabulary, continuous speech recognition (LVCSR) system, spoken input is converted into a string of hypothesized, possibly erroneous, words. However, the current state-of-the-art speech recognition technology is still not robust to all variability in speech signals, especially in a hands-free application. To make the signal pick-up from a speech source more immuned to noise or room reverberation, a microphone array can be employed to form a delay-and-sum fixed beamformer (FBF) toward the sound source or to cancel interferences with a generalized sidelobe canceller (GSC). As a result, the recognition performance in a hands-free operation can be improved but occasional recognition errors are still unavoidable. In this paper we will investigate how to assess the reliability of LVCSR recognized words in a statistically meaningful manner and tag a confidence measure to each recognized word such that an appropriate acceptance/rejection decision can be taken.