Multilingual Document Alignment - A Study with Chinese and Japanese.
Md Maruf Hasan, Yūji Matsumoto · NLPRS · 2001
Natural language processing (NLP) community is increasingly using paralleland comparablecorpora for cross-linguistic research. The knowledge extracted from such corpora helps us in cross-language information retrieval, topic detection and tracking, machine translation, and many other NLP tasks. Parallel or comparable corpora of JapaneseChinese language-pair are rare. We investigate an automatic approach to build bilingual corpora from a collection of unaligned bilingual-documents using linguistic and statistical processing. The similarities between two documents across the languages are calculated using mutual information (MI) and residual inverse document frequency (RIDF) of Kanji. We explained the document alignment algorithm in detail and evaluated the effectiveness of the algorithm using a collection of potentially relevant but unaligned bilingual documents.