Si-Ta: Machine Translation of Sinhala and Tamil Official Documents
Surangika Ranathunga, Fathima Farhath, Uthayasanker Thayasivam, Sanath Jayasena, Gihan Dias · 2018
Although Sri Lanka is a multi-ethnic country with Sinhala and Tamil being the official languages, most of the population is familiar with only one of these languages. This results in a lack of Sinhala-Tamil translators, which in turn has an impact on the government agencies that are required to issue official government documents in both languages. Although Machine Translation can be considered as a possible solution, available translation systems for this language pair have a poor performance, mainly because they do not focus on official government documents. This paper presents Si-Ta, the first Machine Translation system for Sinhala and Tamil official government documents. Si-Ta uses a Statistical Machine Translation engine, and provides a user-friendly web interface. Results show that Si-Ta can be used to eliminate the need for manual translators, if the only requirement is to understand the document received in the source language. In other words, current version of Si-Ta is capable of translating without loss of semantics at a level that is enough for any common target language reader to understand the message in source language.