Developing corpus of Japanese-English Singular Sentence Textual Entailment
Daiki Hayakawa, Masatoshi Tsuchiya, Hitoshi Isahara · 2016
There is an increasing interest against cross lingual textual entailment (CLTE), which is one of the important tasks to improve cross lingual information reliablity, in recent years. This paper describes our ongoing project to develope a Japanese-English CLTE corpus consisting of many singular sentence pairs which are labeled as either entailed or unentailed. This paper proposes (1) using a bilingual parallel corpus as a source of CLTE singular sentence pairs and (2) using the hierachical document structure of the source parallel corpus to select appropriate sentence pairs among vast possible ones. The experiment against Japanese-English Bilingual Corpus of Wikipedia's Kyoto Articles shows that our proposing method is more promising to construct a CLTE corpus than a random sampling method.