Sandhi Splitter for Malayalam Using MBLP Approach

M. Nisha, P. C. Reghu Raj · Procedia Technology · 2016

The morphological richness and the agglutinative nature of Malayalam make it necessary to retrieve the root word from its inflected form in most of the NLP tasks. This paper presents an approach to identify the suffixes of Malayalam words using MBLP approach. The idea here is to use Memory Based Language Processing (MBLP) algorithm for Malayalam suffix identification. MBLP is an approach to language processing based on exemplar storage during learning and analogical reasoning during processing. Sandhi splitting is essential for morphological analysis, document indexing and topic modeling. Suffix separation improves the quality of machine translated text. Training instances created from words are manually annotated for their segmentation and the system is trained using TiMBL (Tilberg Memory Based Learner). The paper presents memory-based model of Malayalam suffix identification and its generalization accuracy.

Read the paper · More papers on PaperTik