Collection and analysis of a Japanese-English emphasized speech corpora
Quoc Truong, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura · 2014
Speech-to-speech (S2S) translation [10] is gradually starting to break down the language barrier, bringing opportunities for people to understand each other using different languages. However, one of the limitations of current S2S systems that they usually do not translate the paralinguistic information included in the input speech. Among the various types of paralinguistic information, we focus on emphasis, a type of information that is used to convey the focus of the sentence, emotion of the speaker, or other high level information useful for communication. This paper describes the collection of an Japanese-English emphasized speech corpora that can be used in the study of how emphasis is expressed across languages. We constructed 2 corpora, one containing digit strings and one with a utterances from a conversational setting. The speakers who can speak both Japanese and English were selected for the recording. 500 parallel digit strings for the digit corpus and 2030 parallel sentences for the conversation corpus were collected. The corpora may be used to analyze emphasis of one language or between languages, or develop emphasized speech translation systems.