A learning word-spotting method for speaker-independent word recognition in noisy environments
Hiroshi Kanazawa, Yoichi Takebayashi · The Journal of the Acoustical Society of America · 1989
A learning word-spotting method has been proposed for the purpose of robust speaker-independent word recognition in noisy environments. In order to avoid word boundary detection errors at the recognition stage, the method employs word spotting based on the multiple similarity method, which was shown to be effective for noisy speech data. The learning process uses synthesized noisy speech data, a mixture of pure speech data and noise data, to design reliable word reference vectors for the word spotting. Word feature vectors with maximum multiple similarity values are automatically extracted by the word spotting. During the learning process, the signal-to-noise ratio (SNR) of the synthesized data is gradually decreased to perform word spotting accurately. Experiments were carried out for 13 words, including ten Japanese digits, spoken by 50 males. Under the condition of 10 dB, SNR contaminated by concourse noise, the recognition scores of 85.5 % and 94.1% were obtained by word spotting, without learning and with learning, respectively. The results have shown the effectiveness of the proposed method in noisy environments.