Rule based Kannada Agama Sandhi splitter
Hosahalli Lakshmaiah Shashirekha, K. S. Vanishree · 2016
The development of tools and techniques for automatic processing of natural language texts at different levels for various applications including machine translation needs to be addressed with at most priority. Sandhi splitter is one such automated tool which acts as a preprocessing tool for morphological analyzer that identifies the morpheme boundaries in a compound word based on the Sandhi rules defined in reverse direction. Due to highly agglutinative and morphologically rich nature of Kannada - one of the Dravidian languages of South India, complex compound words are formed by combining more than one morpheme/word based on the Sandhi rules. In this paper, as a preliminary work we have presented a rule based Kannada Agama Sandhi splitter, which identifies two flavors of Agama Sandhi namely, Yakaragama and Vakaragama and splits the compound word according to the rules of these two Agama Sandhis. Syllables are extracted from the word to identify the split point of the word. Our model uses a dictionary of root words and suffixes to check the validity of split words and rules to identify the split point. The problems of romanization and transliteration are overcome by using Unicode with UTF-8 representation for Kannada. Our model is tested on a large list of words extracted from Prajavani - a popular Kannada daily newspaper and other online resources, and the results are illustrated.