DISTANTSPEECHRECOGNITION:BRIDGINGTHE GAPS JohnMcDonough' 3
Wapler Matthias · 2008
bad,frombothsides. Herewehopetooffer aunique perspective frompeople withonefoot inbothworlds. Whilegreat progress hasbeenmadeinbothfields, there iscur- Inparticular, ourdiscussion isorganized around five gapsinthe rently arelatively large rift between researchers engaged inacousticemerging field ofdistant speech recognition (DSR). Ineachofthese array processing andthose engaged inautomatic speech recognition. gaps, weperceive thepossibility ofmakingsignificant progress in Thisisunfortunate formanyreasons, butmostofall because itpre- thecoming years through achange ofresearch paradigm, sotosay, vents thetwosides, bothofwhomareinvestigating different aspectsbylooking attheproblem through neweyes. Thefirst gapconcerns ofthesameproblem, fromtruly understanding oneanother andco- theunification oftheresearch community involved inconventional operating. Inmanycases, thetwosides seeeachother through the beamforming withthat involved inindependent component analyeyesofstrangers. Ifground breaking progress istobemadeinthe sis(ICA). Eachcommunity examines thesameproblem, butdefines emerging field ofdistant speech recognition (DSR), this abysmalitself bywhatknowledge itdoesnotconsider: Thosedoing convenstate ofaffairs mustchange. Inthis work, weoutline five pressingtional beamforming confine themselves tousing second-order statisproblems intheDSRresearch field, andwemakeinitial proposals tics. Thoseactive intheICAfield consider nogeometric informafortheir solutions. Theproblems discussed herearebynomeans tion. Henceourquestion, whycannot algorithms beformulated that theonly onesthat mustbesolved inorder toconstruct truly effective utilize bothhigher order statistics aswellasgeometric information? DSRsystems. Nonetheless, their solution, inourview, will representItmayarguably besaid that neither source ofinformation issuffisignificant first steps towards this goal, inasmuch asthesolution of cient forbuilding effective DSRsystems. Butperhaps whenused eachofthese problems will require asubstantial change inthemind- together withabitofinnovation, theyaresufficient. sets andthought patterns ofthose engaged inthis field ofresearch. Thesecond gappertains totheformulation ofaconsistent apIndexTerms-speech feature enhancement, particle filter, proach towards combating thetwomostprominent distortions inmulti-step linear prediction, joint denoising anddereveberation, troduced byrealistic acoustic environments, namely, noise andreautomatic speech recognition, beamforming, microphone arrays verberation. Allknowntechniques, suchasspectral subtraction or multi-stage linear prediction, aredesigned tosuppress only oneof thses twodistortions. Werefer tocurrent workbased onthecombi