Automatic Evaluation of Pronounciation in Language Classes
Yoshiyuki Kawasoet, Hiroshi Kanail · 1989
Introduction Due to general trend in international society, to obtain multiple language skill is becoming necessary. Especially, language skill focused on conversation is now thought to be fundamentally important. Up to the present, most frequently used method to learn pronounciation is to listen to recorded materials. Language laboratory is an advanced version of this lesson. These methods, however, are not satisfactory for the evaluation of achievement level of each student. Everywhere there are absolute lack of good trainers. Moreover, even well trained teachers sometimes cannot explain the difference in pronounciation explicitly to their students. Their classical explanation is qualitative and trainees are only forced to follow the model pronounciation. On the other hand, recent progress in speech recognition study is dramatic. We planned to apply the result of the speech recognition study to language education[ 1). Although the system was originally aimed at the Japanese Training Course for foreign students of Tohoku University, it has been designed and developped to fit general purpose. Model pronounciation is presented to trainee. The trainee follows the given example and his/her voice is recorded. The voice is analyzed by linear prediction. The Cepstrum coefficients are numerically deduced and the time scale is altered to fit to the model pronounciation by dynamic programming (DP) method[2]. The DP path accordingly obtained is a measure of global and detailed speed differences of the trainee’s pronounciation. The obtained non-linearly matched result gives a good measure of characteristic features of trainee’s pronounciation. The most important features we have extracted for the evaluation of the voice are (1) timing (speed), (2) power (stress), (3) pitch (accent, intonation), and (4) formants (tone). The next section introduces the principal concept of the present study. Following the method indicated, we have performed experimental application of the developped system. Several typical examples of pronounciations are tested and the extracted features are discussed in the following. Principal Concept of the Present Study Recorded voice is A/D converted with labits accuracy and the sampling rate of 10kHz. Linear prediction is applied to the digitized data with the frame length of 25.6ms and the frame interval of 10ms. After these procedures, the following characteristic features of the voice are extracted.