Document and sentence alignment in comparable corpora using bipartite graph matching

Zeinab Rahimi, Kaveh Taghipour, Shahram Khadivi, Nasim Afhami · 2012

Parallel corpora are considered as an inevitable resource of statistical machine translation systems, and can be obtained from parallel, comparable or non-parallel documents. Parallel documents are more suitable resources but due to their shortage, comparable and nonparallel documents are also used. In this paper, we address both document alignment and sentence alignment in comparable documents as an assignment problem of bipartite graph matching and intend to find the sub graphs having the maximum weight. One of the best methods to solve this problem is Hungarian algorithm which is a combinatorial optimization problem with known mathematical solutions. The advantages of proposed method are language independency and time complexity of O(n3) for Hungarian algorithm. We have applied this method to bilingual Farsi-English corpus, and obtained high precision and recall for this method.

Read the paper · More papers on PaperTik