Japanese dictation system using character source modeling
Tatsuro Yamada, SHINYA MATSUNAGA, Kiyohiro Shikano · 1992
The authors describe a Japanese dictation system that uses a stochastic language model based on sequences of Japanese characters. The trigram probabilities, which are obtained from a text database consisting of Kanji and Kana are used to construct a source model. A Japanese dictation system generally requires Kana-to-Kanji conversion if the system uses a phoneme based unit for the acoustic processing. However, a system that uses a Kanji-and-Kana character source model can generate an output Kanji-and-Kana sequence directly from input speech without using Kana-to-Kanji conversion. The system is tested using 274 phrases uttered by one male speaker, and achieves 58.4% phrase transcription rate. When the system uses a pronunciation dictionary and eliminates the candidates whose Kanji readings are contextually inappropriate, the phrase transcription rate increases to 63.9%. It is confirmed that a Japanese character source model is efficient for a Japanese dictation system.>