Word-spotting based on inter-word and intra-word diphone models
T. Nitta, Shinichi S. Tanaka, Yasuyuki Masai, Hiroshi Matsuura · 2002
The authors propose a precise but simple inter-word diphone model (IDM) for word-spotting based on SMQ/HMM. They have applied ordinary diphone models to a speaker-independent, large-vocabulary word recognition unit. However, because users are apt to add words and/or extraneous speech, accuracy degrades due to the mismatch of models at word-boundaries. The IDM represents a transition from the preceding phonemes to a word or from a word to the succeeding phonemes. An experiment showed that the IDMs reduce error rates by about 5% for speech containing unknown words and extraneous speech. The experiment also showed that the proposed method ensured performance good enough for the practical use of a large-vocabulary isolated-word recognition system.