Acoustic analysis and modeling of speech based on phonetic features

Carol Espy-Wilson, Nabil N. Bitar · 1998

Acoustic modeling and analysis of speech based on phonetic features is explored in the current research for speaker-independent speech recognition. Phonetic features are minimal speech units that describe the manner and place of articulation of the sounds of a language. In this research, it is shown that phonetic features have acoustic signatures in the speech signal that can be reliably extracted in a manner that reduces the effects of speaker-differences. Moreover, it is postulated based on the conducted experiments that using phonetic features as the basic speech units allows for the modeling of contextual variability in a general and natural way. A major thrust of this thesis is in the development of algorithms that extract acoustic properties of the phonetic features. These algorithms make measurements on the speech signal that are motivated by acoustic phonetics and spectrographic analysis. A measurement is made at a time-instant relative to its value at another instant and/or is made in a frequency band relative to another. Such relative measurements focus on the linguistic content of the speech signal reducing the effects of interspeaker variability. In one part of this thesis, acoustic measurements were developed based on subjective acoustic analysis. An event-based recognition system that uses these measurements, combined by fuzzy rules, was developed and compared to a Hidden Markov Model (HMM) system using (1) the same measurements but modified to fit the frame-based HMM system and (2) Mel-cepstral parameters. The results show that the event-based approach produces comparable results to the HMM frame-based system for the undertaken task of broad-class speech recognition. In addition, it is shown that the developed measurements perform better than the cepstral parameters in this task. An automatic optimization procedure based on the Fisher criterion and classification trees was developed to automate the derivation of acoustic measurements. Using this procedure, manner and place-of-articulation acoustic measurements were developed. These measurements were evaluated in phonetic-feature classification tasks and in a 10-class recognition task using an HMM system. Recognition results compared favorably to those obtained with Mel-cepstral parameters. The results show that the developed measurements target the intended linguistic information and are robust to speaker differences.

Read the paper · More papers on PaperTik