Data Collection and Transliteration of Japanese Spontaneous Database in the Travel Arrangement Task Domain

Akira Kurematsu, Youichi Akegami, Schultz, Tanja, Burger, Susanne · 1999

This paper describes the method to construct and transcribe Japanese spontaneous speech data for VERBMOBIL, the German research project of speech translation.. Spontaneous spoken dialogue database is the basis for developing speech and language processing for dialogue systems such as speech translation system. The extended data of human-to-human spoken dialogue in the scenario of travel arrangement has been initiated to be collected in German, English and Japanese in the travel arrangement task. Romanized transcription is used to develop acoustic model and language model in speech recognition system, and natural language translation system. In this paper, issues of transliteration method and several rules and conventions to transcribe Japanese spoken dialogue will be described.

Read the paper · More papers on PaperTik