Design and Development of Japanese Speech Corpus for Large Vocabulary Continuous Speech Recognition Assessment

Katsunobu Itou, Kazuya Takeda, Toshiyuki Takezawa, Tatsuo Matsuoka, Kiyohiro Shikano, Tetsunori Kobayashi, Shuichi Itahashi, Mikio Yamamoto · Institutional Repositories DataBase (IRDB) · 1998

The project of developing LVCSR (Large Vocabulary Continuous Speech Recognition) platform is introduced. It is a collaboration of researchers of different academic institutes and intended to develop a sharable software repository of not only databases but also models and programs. The platform consists of a standard recognition engine, Japanese phone models and Japanese statistical language models. Japanese acoustic-phonetic models are trained with ASJ (Acoustic Society of Japan) databases and scalable from a context-independent monophone to a triphone model of thousands of states. Japanese word N-gram (2-gram and 3-gram) models are constructed with a corpus of Mainichi newspaper of four years. The recognition engine JULlUS is developed for assessment of both acoustic and language models. As an integrated system of these modules, we have implemented a baseline 5000-word dictation system and evaluated various components. The software repository is available to the public.

Read the paper · More papers on PaperTik