Research and development of robust speech recognition

Kiyoaki Aikawa · The Journal of the Acoustical Society of America · 1998

This paper describes recent research and development activities on robust ASR (automatic speech recognition) in NTT Human Interface Laboratories. ASR system design has been changing from the experimental to the commercial level. A relevant issue in achieving practical ASR is the robustness against environmental noise, and speaker and circuit differences. Adaptation technique has been widely investigated for improving the robustness. It should be noted that customers or users generally prefer simple procedures and quicker response. Thus real time speaker/noise adaptation is getting more and more popular. Confidence measure is also effective in enhancing speech recognition performance by rejecting out-of-vocabulary words, lip, and breath noises. Another relevant issue is how to quickly reach the goal of the human-machine dialog via the telephone without visual information. The voice quality through a telephone is not good enough to provide perfect syllable intelligibility even for a human. This suggests that an appropriate human-machine dialog is needed to support an ASR system in acquiring the information or request intended by the user. Recent research topics including spectral estimation, dialog control, and other new approaches related to above discussion will be shown in the presentation. ASR application examples and related problems will also be shown.

Read the paper · More papers on PaperTik