Compiling a Turkish-English Bilingual Corpus and Developing an Algorithm for Sentence Alignment
Alper Güngör, Tunga Güngör · 2006
Abstract: In this paper, we discuss the compilation of a bilingual Turkish-English corpus and propose a method for sentence alignment based on location and length information in the texts. The content of the corpus was collected from several sources of different genre and it contains about 5 million words. To the best of our knowledge, this is the first comprehensive bilingual corpus between these languages. The proposed sentence alignment algorithm was tested on the corpus and success rates up to 96 % were obtained.