Time-varying feature selection and classification of rapidly changing speech signals
K. S. Nathan · 1992
A parametric time-varying model is developed and is used to generate a feature set that captures the dynamics of rapidly changing speech signals. The performance of this time-varying feature set on such signals, e.g. transitions into and out of stop consonants, is studied. Although several such models have been proposed in the recent literature, most employ least-squares methods for parameter extraction which do not perform very well in the presence of additive noise or when the magnitudes of some of the coefficients are very small, both distinct possibilities when dealing with real speech data. Furthermore, these models fail to differentiate regions of glottal excitation from those where the glottis is closed. The elimination of glottal effects can lead to significantly more accurate parameter estimates. In this work, each of these issues is addressed. An all-pole model with linearly time-varying predictor coefficients is the basis for the system. Only data corresponding to the closed-glottis portion of a pitch period are considered. Accordingly, a procedure to identify and extract regions of the data which best fit the model is developed. An efficient, iterative algorithm solves the non-linear optimization problem that results from the maximum likelihood estimate of the parameters. The resulting parameters are used to factorize the time-varying formant trajectories over the interval of analysis. A general feature set consisting of the time-varying predictor coefficients and formant trajectories is developed. The time-varying feature set is utilized to classify the unvoiced stop consonants. A novel adaptive classifier, Learning Vector Classifier, (LVC), is also described, and its performance is compared to that of other similar classifiers. Very high, speaker-independent, recognition rates, $\approx$95%, are obtained using no information from the release of the consonant. Only data concerning the second formant prior to closure are used. The high recognition rates are evidence of the effectiveness of this time-varying feature set. Further improvements in recognition accuracy are anticipated when information from the burst in conjunction with a more complete time-varying feature set is used.