The Kazakh-Russian parallel corpus of criminal texts

Orken J. Mamyrbayev, Nina Khairova, Waldemar Wójcik · 2025

Linguistic resources are crucial for both linguistic research and the development of natural language processing applications, such as machine translation and text summarization. Among these, the creation of high-quality text corpora is a vital area of research, enabling detailed statistical analysis and exploration of linguistic changes over time. Among the various categories of the corpora, the parallel corpora play a significant role in studying language features and improving machine translation quality. This chapter shows the developed corpus of Kazakh and Russian criminally related texts and describes the corpus’s structure and context. Four bilingual websites were selected for text collection: zakon.kz, caravan.kz, lenta.kz, and nur.kz, all of which are well-known and reliable portals in the Republic of Kazakhstan, with criminal news being one of their primary categories. The performed expert evaluation showed that the accuracy of automatically aligned sentences of the created parallel Kazakh-Russian corpus is about 60%, with a coefficient of agreement of 0.83. The other sentences were aligned manually.

Read the paper · More papers on PaperTik