English-Kazakh Parallel Corpus For Statistical Machine Translation

Ayana Kuandykova, Amandyk Zhankozhauly Kartbayev, Таннур Ерланович Калдыбеков · International Journal on Natural Language Computing · 2014

This paper presents problems and solutions in developing English-Kazakh parallel corpus at the School of Mechanics and Mathematics of the al-Farabi Kazakh National University.The research project included constructing a 1,000,000-word English-Kazakh parallel corpus of legal texts, developing an English-Kazakh translation memory of legal texts from the corpus and building a statistical machine translation system.The project aims at collecting more than ten million words.The paper further elaborates on the procedures followed to construct the corpus and develop the other products of the research project.Methods used for collecting data and the results are discussed, errors during the process of collecting data and how to handle these errors will be described.

Read the paper · More papers on PaperTik