A novel approach to Sandhi splitting at character level for Kannada language

M. Rajani Shree, Sowmya Lakshmi, B R Shambhavi · 2016

Natural Language Processing (NLP) is a field of computational linguistics related to interactions between computers and human languages. Parsing of input text of any human language is a part of NLP. This parsing technique requires processing of text on a word by word basis. To process any individual word especially in Sanskrit and all Dravidian languages, Sandhi splitting is a major task. Sandhi is also called Morphophonemics concerned with changes that occur when two words or separate morphemes come together to form a new word. Exact splitting point is essential for text processing tasks such as POS tagging and in turn parsing. We have adopted a novel approach to internal Sandhi splitting technique on Kannada language. Each Kannada word is split into morphemes according to valid morph patterns. After the division of each word into lexical morphemes, we have manually tagged each split word into root-begins, root-continuous and suffix. This work has been done on Kannada language. We have trained the system with a list of 1000 tagged words using a CRF (Conditional Random Fields) tool and nearly 400 raw split words (untagged words) are given as input to the CRF tool. The system generates a list of tagged split words for the given input according to the trained data. The system output has been compared with the manually tagged data. We have verified the data using 5 fold test, which takes five different combinations of trained data (1000 words) and test data (400 words). The average Precision, Recall and F-Measure of Tagging accuracy of CRF model for Kannada corpus in 5 fold test are nearly equal to 98.08, 92.91 and 95.43. This method can be successfully implemented in all other Dravidian languages for the Sandhi splitting.

Read the paper · More papers on PaperTik