Compiling a Turkish-English Bilingual Corpus and Developing an Algorithm for Sentence Alignment

Alper Güngör, Tunga Güngör · 2006

Abstract: In this paper, we discuss the compilation of a bilingual Turkish-English corpus and propose a method for sentence alignment based on location and length information in the texts. The content of the corpus was collected from several sources of different genre and it contains about 5 million words. To the best of our knowledge, this is the first comprehensive bilingual corpus between these languages. The proposed sentence alignment algorithm was tested on the corpus and success rates up to 96 % were obtained.

Read the paper · More papers on PaperTik