Continuous multiple similarity method using stable and dynamic phonetic feature vectors for continuous speech recognition
H. Tsuboi, Yoichi Sadamoto, Yoichi Takebayashi · The Journal of the Acoustical Society of America · 1990
This paper describes the continuous multiple similarity method (CMSM) that utilizes stable and dynamic phoentic feature vectors to achieve continuous speech recognition. The phonetic feature vector to represent the time-varying characteristics of both stable and dynamic segments is constructed by a fixed dimensional time-frequency spectrum. Stable phonetic segments correspond to the stable portion of five vowels, fricatives and nasals. Dynamic phonetic segments correspond to the time-variant portion of stops, semivowels, and liquids. The stable phonetic segments are represented by a time-series of the time-frequency spectrum, whereas the dynamic phonetic segments are represented by a single typical time-frequency spectrum. The multiple similarity (MS) values corresponding to the particular phone class are time-continuously computed. The time sequence of the MS values are then used for word matching with word graphs by dynamic programming. An experiment was carried out on 6596 phonetic segments of 492-word vocabulary spoken by four Japanese males to evaluate phone recognition performance based on the CMSM. A recognition score of 73.5% for 23 phone classes was obtained using 224-dimensional phonetic feature vectors.