A study on hands-free speech/speaker recognition

Longbiao Wang · Institutional Repositories DataBase (IRDB) · 2008

Automatic speech recognition (ASR) systems are known to perform reasonably well when the speech signals are captured using a close-talking microphone. However, there are many environments where the use of a close-talking microphone is undesirable for reasons of safety or convenience. Hands-free speech communication has been more and more popular in some special environments such as an office or the cabin of a car. In this thesis, robust speaker localization, speech recognition and speaker identification methods under a distant-talking environment (that is, a hands-free environment) are presented. In this thesis, we first propose a robust speech/speaker recognition method by incorporating the estimated speaker position information into Cepstral Mean Normalization (CMN), which is called Position-Dependent CMN (PDCMN). The system measures the transmission characteristics (the compensation parameters for position-dependent CMN) from some grid points in the room a priori. Four microphones are arranged in a T-shape on a plane, and the sound source position is estimated by Time Delay of Arrival (TDOA) among the microphones using a proposed closed-form solution. The system then adopts the compensation parameter corresponding to the estimated position and applies a channel distortion compensation method to the speech (that is, position-dependent CMN) and performs speech/speaker recognition. For close-talking speaker recognition, a speaker recognition method by combining speaker-specific GMMs with speaker-adapted syllable-based HMMs has been proposed, which shows robustness for the change of the speaking style in a close-talking environment. In this thesis, we extend this combination method to distant speaker recognition and integrate this method with the proposed position-dependent CMN. Speaker recognition experiments were conducted on NTT database (22 male speakers) and Tohoku University and Panasonic database (20 male speakers). The speaker identification result by GMM showed that the proposed position-dependent CMN achieved the relative error reduction rates of 64.0% from w/o CMN and 30.2% from the position-independent CMN with the NTT database. By integrating the position-dependent CMN into the combination use of speaker-specific GMMs and speaker-adapted syllable-based HMMs, a furthermore improvement was obtained. For speech recognition, we propose a robust speech recognition by combining multiple microphone-array processing with the position-dependent CMN. In a distant environment, the speech signal received by a microphone is affected by the microphone position, the distance and direction from the sound source to the

Read the paper · More papers on PaperTik