Improving Automated Alignment in Multilingual Corpora

John A. Campbell, Niladri Chatterjee, Mauro Manela, Alex Chengyu Fang · Institutional Repositories DataBase (IRDB) · 1996

We report on methods of improving multilingual text alignments that have been produced in a simple dynamic-programming scheme, by automated detection of possible misalignments.Details of methods involving cognates, speciallyidentified words, and propositional contents of sentences are given, together with notable features of their performance on parallel corpora in a number of different types of European languages.Frequency Method English-French English-Czech French-Czech -U 95.7-0.2 95.6-0.2 95.6-0.2 -L 95.5-0.3 95.5-0.2 95.4-0.2 -S 95.6-0.3 95.5-0.2 95.5-0.2 -U+L 95.9-0.3 95.8-0.2 95.8-0.2 -U+S 95.9-0.3 95.9-0.2 95.8-0.2 2 L 95.3-0.2 95.3-0.3 95.4-0.3 2 U+L 95.9-0.4 95.8-0.3 95.8-0.3 4 L 95.7-0.2 95.7-0.2 95.8-0.1 4 U+L 96.2-0.2 96.1-0.2 96.1-0.2 6 L 95.3-0.1 95.3-0.0 95.1-0.1

Read the paper · More papers on PaperTik