Improved vowel recognition using the discrete cosine transform.
Dc Burr · The Journal of the Acoustical Society of America · 1992
Some experiments are described that use a discrete cosine transform (DCT) to represent vowel spectra for classification by a neural network. The recognition accuracy of a neural network trained by back propagation was compared to that of a simple Gaussian classifier. Twelve English vowel categories from the TIMIT database were used {aa, ae, ah, ao, ax, eh, er, ih, ix, iy, uh, uw}. Training was done in a speaker-dependent manner using 152 male speakers from all geographic regions. Eight 16-ms DFT spectra were computed at 8-ms intervals about the center of each vowel. Sixteen DCT coefficients were derived from the 250- to 4250-Hz interval in each 8-kHz spectrum. Average DCT and delta DCT vectors over the eight frames were used as input features. Additional features included the first six peak frequencies of the DCT spectrum and two pitch parameters from the DCT coefficients. Best performance was obtained by the neural network with all 40 input features, resulting in 58.2% recognition accuracy. This compares favorably to a cochleagram representation using the same vowel classes [Muthusamy etal., ICASSP-90]. The DCT also appears to require fewer coefficients than an equivalent cepstrum-based vowel classifier [Burr, ICASSP-92].