Adaptive beamforming and soft missing data decoding for robust speech recognition in reverberant environments
Marco Kühne, Roberto B. Togneri, Sven Erik Nordholm · 2008
Abstract This paper presents a novel approach to combine microphonearray processing and robust speech recognition for reverberantmulti-speaker environments. Spatial cues are extracted from amicrophone array andautomaticallyclusteredtoestimate local-ization masks in the time-frequency domain. The localizationmasks are then used to blindly design adaptive filters in order toenhancethesourcesignalspriortomissingdataspeechrecogni-tion. A novel evidence model better exploiting the informationprovided by the source separation stage is proposed. Recogni-tion experiments demonstrate the effectiveness of the schemewhen compared to traditional microphone array enhancementand a related binaural separation model. Index Terms : missing data speech recognition, microphone ar-ray processing, adaptive beamforming 1. Introduction A key requirement for automatic speech recognition (ASR)technology to be employed in everyday situations is robust-ness to multiple competing speakers in reverberant enclosures.While today’s ASR systems still fall short in comparison withhuman listeners some promising success has been achieved indealing with multiple speakers in anechoic conditions. For ex-ample, the work reported in [1] utilizes binaural localizationcuessuchasinterauraltimeandintensitydifferencestoestimatetime-frequency (TF) masks for missing data speech recognition(MD-ASR).However,inreverberantenclosuresthelocalizationcues become increasingly unreliable making an accurate esti-mation of the TF mask considerably more difficult. Further-more when the speech models are trained on anechoic data thespectral features employed in MD-ASR are adversely affectedby the room reflections.Previous attempts to handle multisource reverberant en-vironments include dereverberation filters [2], adaptation ofspeech models to the room [3], an inhibition mechanism to em-phasizesoundonsets[4]andfeatureenhancementusinganane-choic speech prior based reconstruction technique [5].This paper proposes an alternative approach by extend-ing the two-channel system developed in [6] to the multi-microphone case (see Figure 1a). The system consists of asource separation stage coupled with a missing data decoderthrough the probabilistic concept of evidence models [7]. Forthe source separation stage a fuzzy clustering approach is em-ployed using direction of arrival (DOA) values of the two outersensors in the array. The clustering produces a source DOAestimate and a TF mask marking the dominant TF points foreach source. While in [6] the correlation between adjacent TFpoints wasonly utilized duringmask post-processing inthispa-per the neighborhood information is integrated during the clus-tering process itself. This biases the solution towards homoge-nous masks and helps greatly to reduce the effect of noise visi-ble in the localization cues under reverberant conditions.Motivated by [8] the masks and source DOAs are then uti-lized to blindly design multiple adaptive beamformers capableofenhancingthespeechsourcespriortorecognition. Thisworkproposes a novel evidence model incorporating both localiza-tion information and feature uncertainty. The system operatesinanunsupervised manneranddependsneither ontrainingdatafor source localization nor a priori learned echoic speech mod-els adapted to the reverberant environment.The remainder of this paper is as follows: Section 2 de-scribes the source separation stage in more detail. Section 3discusses the missing data recognizer and the evidence modelparameter estimation. Section 4 reports the evaluation resultsand compares these with a related binaural model. The papercloses in Section 5 with an outlook into future work.