Evaluation of ASR front ends in speaker-dependent and speaker-independent recognition

Jean-Claude Junqua · The Journal of the Acoustical Society of America · 1987

This paper extends previous experiments of Tsuga and Hermansky [J. Acoust. Soc. Am. Suppl. 1 80, Sl8 (1986)]. Those experiments evaluated the effect of spectral model order, in automatic speech recognition (ASR), using a small alpha-numeric data base. PLP (perceptually based linear predictive) and LP (linear predictive) analyses were compared, using a cepstral and RPS (root power sums) metric. Those experiments dealing with a bigger data base were validated (104 words, ten speakers). PLP RPS front end is compared with about ten other ASR front ends (LP ceptrum, LP RPS, critical band,…). Experiments were run at various different analysis model orders. Results of speaker-dependent ASR show that the low-dimensional PLP analysis is a good alternative to high-dimensional LP or filter bank analysis. The index-weighted metric improves the recognition accuracy, making the recognition results more uniform across the speakers. The speaker-independent experiments (which use the templates of one male and one female speaker as references) confirm the superior performance of the PLP method for extracting speaker-independent information. Most of the errors are on the consonants; some improvements on the recognition process are being investigated.

Read the paper · More papers on PaperTik