A Rule Based Approach for Root Word Identification in Malayalam Language
Meera Subhash · International Journal of Computer Science and Information Technology · 2012
Words are tools of life which is omnipresent in every language.All words in a language are unique having their own function and meaning.The syntactic and semantic knowledge about individual words can be encapsulated in a highly structured repository known as computational lexicon which is very essential for Machine Translation.For designing a computational lexicon, the first and foremost task is to identify the head words or root words in the language.The Root Word Identifier proposed in this work is a rule based approach which automatically removes the inflected part and derive the root words using morphophonemic rules.The system is tested with 2400 words from a Malayalam corpus to generate the linguistic information such as the root form, their inflected forms and grammatical category.The performance is evaluated using the statistical measures like Precision, Recall and F-measure.The values obtained for these measures are more than 90%.