Mining Patents for Parallel Corpora
Masao Utiyama, Hitoshi Isahara · The MIT Press eBooks · 2008
Large-scale parallel corpora are indispensable language resources for machine translation (MT). However, there are only a few publicly available large-scale parallel corpora. This chapter describes a Japanese-English patent parallel corpus created from patent families filed in Japan and the United States. The parallel corpus contains about 2 million sentence pairs that were aligned automatically. This is the largest Japanese-English parallel corpus and will be available to the public after the NTCIR-7 workshop meeting.