Extraction of fixed dimension patterns from varying duration segments of consonant-vowel utterances
Suryakanth V. Gangashetty, C.C. Sekhar, B. Yegnanarayana · 2004
Classification models based on multilayer perceptron (MLP) or support vector machine (SVM) have been commonly used for complex pattern classification tasks. These models are suitable for classification of fixed di-mension patterns. However, durations of consonant-vowel (CV) utterances vary not only for different classes, but also for a particular CV class. It is necessary t o de-velop a method for representing the CV utterances by patterns of fixed dimension. For CV utterances, vowel onset point (VOP) is the instant at which the consonant part ends and the vowel par t begins. Important infor-mation necessary for classification of CV utterances is present in the region around the VOP. A segment of fixed duration around the V O P can be processed t o ex-tract a pattern of fixed dimension t o represent a C V utterance. Accurate detection of vowel onset points is important for recognition of CV utterances of speech. In this paper, we propose an approach for detection of VOP, based on dynamic t ime alignment between a ref-erence pattern of a CV class and the pattern of an utter-ance of t ha t class. The results of studies show tha t t he hypothesised VOPs using the proposed approach have less deviation from their actual locations. 1.