Analysis of a large in-car speech corpus and its application to the multimodel ASR

Hiroshi Fujimura, Chiyomi Miyajima, K. Itou, Kazuya Takeda, Fumitada Itakura · 2006

In-car ASR performance improvement, utilizing a large in-car speech corpus, consisting of the utterances of more than five hundred drivers under real driving conditions, is discussed. A subset design method for efficient cross validations in large-scale speech recognition experiments is proposed. The factor analysis of the results of the recognition experiments show the relationship between word accuracy and utterance characteristics, i.e., SNR, entropy and speaking rates. Based on the factor analysis results, a multimodel approach which uses the utterance duration and subband SNRs as the model selection measures for acoustic and language models, respectively, is proposed. By the proposed multimodel approach, a relative error reduction of 16% is obtained.

Read the paper · More papers on PaperTik