Voted NER system using appropriate unlabeled data

Asif Ekbal, Sivaji Bandyopadhyay · 2009

This paper reports a voted Named Entity Recognition (NER) system with the use of appropriate unlabeled data.The proposed method is based on the classifiers such as Maximum Entropy (ME), Conditional Random Field (CRF) and Support Vector Machine (SVM) and has been tested for Bengali.The system makes use of the language independent features in the form of different contextual and orthographic word level features along with the language dependent features extracted from the Part of Speech (POS) tagger and gazetteers.Context patterns generated from the unlabeled data using an active learning method have been used as the features in each of the classifiers.A semi-supervised method has been used to describe the measures to automatically select effective documents and sentences from unlabeled data.Finally, the models have been combined together into a final system by weighted voting technique.Experimental results show the effectiveness of the proposed approach with the overall Recall, Precision, and F-Score values of 93.81%, 92.18% and 92.98%, respectively.We have shown how the language dependent features can improve the system performance.

Read the paper · More papers on PaperTik