CRF Based POS Tagging of Manipuri

Victor Haobam, Laishram Jimmy, Nirman Longjam · 2010

Abstract This paper deals about the tagging of Part of Speech (POS) in the Manipuri text using Conditional Random Field (CRF) which is an unsupervised learning approach. Manipuri is a language spoken mainly in Manipur, a state in the North Eastern Part of India also in some other part of Myanmar and Bangladesh. It is highly agglutinative in nature thus choosing of feature such as surrounding words, POS tag, prefix, suffix, length etc as feature for running CRF tool is helpful for POS tagging. The system shows a recall of 70.00%, precision of 77.78% and F-measure of 73.68%. Keywords: Part of Speech (POS), Conditional Random Field (CRF), Feature, Multiword Expression (MWE), Recall, Precision and F-measure. 1 Introduction This paper deals with the unsupervised tagging of Part of Speech (POS) for Manipuri language (or Meiteilon), which is one of the scheduled languages of India mainly spoken in Manipur and in some parts of Assam which locates at the North-eastern part of India also in other countries like Bangladesh and Myanmar. Manipuri uses two scripts; the first one is purely of its own origins while another one is a borrowed Bengali script. In the present task, the processing has been done on the Bengali script. Manipuri is a Tibeto-Burman (TB) language which is highly agglutinative in nature, mono-syllabic, influenced and enriched by the Indo-Aryan languages of Sanskrit origin and English. The affixes play the most important role in the structure of the language. A clear-cut demarcation between morphology and syntax is not possible in this language. In Manipuri, words are formed in three processes called affixation, derivation and compounding as mentioned in Thoudam P.C. The majority of the roots found in the language are bound and the affixes are the determining factor

Read the paper · More papers on PaperTik