On the application of spectrum target prediction model to speech recognition
Masato Akagi, Yoh’ichi Tohkura · 2003
A preprocessing method is proposed for automatic speech recognition that uses a spectrum target prediction model to cope with coarticulation, one of the most serious problems in automatic speech recognition. The method is evaluated by three measures: spectral stability with respect to measuring predicted spectrum variation, and intracategory variation. Experimental results indicate that predicted spectra throughout the model are stabilized in each phoneme portion by eliminating variations of original spectra without prediction. The results also indicate that by using the preprocessing method, intracategory variation decreases and intercategory variation increases. Consequently, the spectrum target prediction model implemented as a speech-recognition preprocessor improves automatic speech recognition performance.>