A language-independent method for the alignement of parallel corpora
Thị Minh Huyền Nguyễn, Mathias Rossignol · Institutional Repositories DataBase (IRDB) · 2006
The automatic alignment of parallel corpora is a very rich source of information for automatic translation, multilingual document indexing, information retrieval, etc. The rapid growth of the use of “ minority ” languages in online documents makes it necessary to develop methods that can easily adapt to any language. We present an evolution over previous works, notably by Church and Gale [1], that performs the alignement of parallel texts in any language without any need for information concerning these languages, as is often the case in existing systems. We also introduce a more experimental system, that shows promise for the alignment of degraded translations. The systems have taken part to the ARCADE II evaluation campaign [2], of which we present the results.