Recognition of vowels in continuous speech based on the dynamic characteristics of WLR distance trajectories.
Yutaka Kobayashi, Yasuhiro Ohmori, Yasuhisa Niimi · Journal of the Acoustical Society of Japan (E) · 1986
In this paper the authors propose a method of vowel recognition in continuous speech and give some experimental results. The dynamic characteristics of speech are analyzed in order to enhance the intended vowels and to prune spurious ones. Physical parameters of vowels in continuous speech hardly reach at their preset target values because of the smoothing effects of coarticulation. However, the direction of the targets are detectable in many cases. The analysis algorithm uses the temporal movements of the Weighted Likelihood Ratios between the input speech and 6 templates: 5 Japanese vowels and a nasal group. Recognition experiments were carried out for two sets of speech data spoken by two male speakers. The sets contain 35 and 53 sentences, respectively. Using the speaker-dependent templates, 76.2 % and 80.7 % of vowels were correctly recognized and the effectiveness of the enhancement algorithm was proved. Major problems left for further improvement are treatments of long vowels, diphthongs, semi-vowels, devocalization, and nasalization.