Creating a Noisy Parallel Corpus from Newswire Articles Using Cross-Language Information Retrieval

Nigel Collier, Hideki Hirakawa, Akira Kumano · 1998

this paper we present an adaptation of cross-language information retrieval for the production of an aligned bilingual corpus from noisy-parallel English-Japanese newswire articles. We implement the standard vector space model and show though simulation the effectiveness of five variations for the alignment task. The methods are computationally efficient, easy to evaluate, and generalizable to other genres and language pairs --- an important factor if we are to use the aligned articles for knowledge acquisition in unrestricted domains. Our results show that alignment precision levels of over 70% at 70% recall are possible. 1. Introduction

Read the paper · More papers on PaperTik