Improved hands-free automatic speech recognition in reverberant environment condition
Randy Gómez, Keisuke Nakamura, Takeshi Mizumoto, Kazuhiro Nakadai · 2014
In this paper, we incorporate language information in the treatment of reverberation in a microphone array system. First, microphone array processing is used to spatially filter the observed multiple sources. Second, the separated signal is enhanced to remove the effects of reverberation. Then, we extract information concerning the impact of smearing to the language (word sequence) used in the automatic speech recognition (ASR) system. Consequently, we compute the language-inferred coefficients via offline training. During online testing, the language-inferred coefficients are employed in conjunction with an enhancement scheme. The coefficients reflect the realistic impact of smearing in conjunction with the language information. As a result, the enhanced signal is more effective in ASR application. Experimental results in a real environment condition show that the proposed method outperforms the conventional methods that exclude language information. Moreover, we show the recognition results comparing the proposed method against existing dereverberation methods.