Effect of the spectral model order in automatic speech recognition

Kazuhiro Tsuga, Hynek Heřmanský · The Journal of the Acoustical Society of America · 1986

It has been observed that in speaker-independent multi-template digit recognition, the 5th-order perceptually based LP (PLP) analysis method yields about 40% lower error rates than does the standard 14th-order LP. In speaker-dependent recognition, the 5th-order PLP model yields essentially the same recognition accuracy as do high-order PLP or standard LP models. In order to clarify the reasons for this result, RPS distances for LP and PLP models of single-frame phoneme-like spectra of male and female speakers were computed and arranged into distance matrices. Single-speaker 10th-order model matrices were adopted as the reference matrices for both LP and PLP methods. Cross-speaker PLP matrices become more similar to PLP reference matrices as the model order of cross-speaker matrices decreases to the 4th order. LP cross-speaker matrices remain relatively similar to the 5th order. Single-speaker PLP matrices remain relatively similar to the 5th order of the model while the single-speaker LP matrices of lower order differ significantly from the high-order ones. These results agree well with recognition rates obtained in the cross-speaker alpha-digit recognition for the same speakers. Results support the F1, F2′ theory of speech perception.

Read the paper · More papers on PaperTik