Automatic detection and modeling of new words in a large-vocabulary continuous speech recognition system
Ayman Asadi · 1992
In practical large vocabulary continuous speech recognition systems, it is nearly impossible for a speaker to remember which words are in the vocabulary. In this thesis, we describe a novel technique that automatically detects when the speaker has used a word that is not in the vocabulary. The technique uses a general model for the acoustics of any word to recognize the existence of new words. Using this general word model, we measure the detection rate of new words versus the false alarm rate. Experiments were run using the DARPA 1000-word Resource Management Corpus for continuous speech recognition. We have used BYBLOS, the BBN continuous speech recognition system, which is based on hidden Markov models, to perform these experiments. The results indicated a useful detection rate for new words of 71% with a false alarm rate of 1%, which imply that the new-word detection technology described in this thesis, can be used in current speech recognition systems. Once a new word is detected, it is desirable to add the word to the vocabulary of the system. We present a new technique for obtaining a phonetic transcription for a new word, which is needed to add the new word to the system. The technique utilizes DECtalk's text-to-sound rules to obtain an initial phonetic transcription for the new word. Since these text-to-sound rules are imperfect, we use a probabilistic transformation technique that produces a phonetic pronunciation network of all possible pronunciations given DECtalk's transcription. The network is used to constrain a phonetic recognition process that results in an improved phonetic transcription for the new word. The improved phonetic transcription is used to build a specific model for the new word. Then the word is added to the vocabulary and to the proper places in the language model.