Estimation of statistical phoneme center considering phonemic environments
Shigeki Okawa, K. Shirai · 2002
The paper presents a new scheme of acoustic modeling for speech recognition based on an idea of statistical phoneme center. The statistical phoneme center has several properties that are feasible for realizing more reliable phoneme extraction. First, the authors assume that there is a fictitious center point in every phoneme. The center is determined statistically by an iterative procedure to maximize the local likelihood using a large amount of speech data. Next, in order to evaluate the performance of phoneme extraction, phoneme recognition is realized by optimizing the likelihood based on the dynamic time warping technique. As an experimental result, 71.6% recognition accuracy is obtained for speaker independent phoneme recognition. This result demonstrate that the proposed SPC is a new effective concept to obtain more stabilized acoustic model for speaker independent speech recognition.