Aligning Sentences for Kazakh-Turkish Parallel Corpora
Diana Rakhimova, Eşref Adalı, Aidana Karibayeva · 2024
This paper presents a hybrid approach to sentence alignment for the Kazakh-Turkish parallel corpus, addressing the challenges posed by linguistic and structural differences between the two languages. The system is divided into two modules: the Hunaling module and the Rule-Based module. The Hunaling module first performs an initial alignment using sentence length to establish a rough correspondence between sentences. This is followed by a refinement stage that leverages lexical matching, focusing on Turkic cognates and shared roots. The Rule-Based module then applies syntactic and contextual rules to fine-tune the alignment, using a small annotated corpus to guide adjustments. The proposed approach achieves a precision of 90%, recall of 88%, and an F1-score of 88.5%, demonstrating significant improvements over traditional methods. This methodology offers a robust solution for aligning sentences in Kazakh-Turkish parallel corpora, with potential applications in machine translation and multilingual NLP.