A Comparative Study of Named Entity Recognition for Hindi Using Sequential Learning Algorithms
Awaghad Ashish Krishnarao, Himanshu Gahlot, Amit Srinet, Dharmender Singh Kushwaha · 2009
Through this paper we present a comparative study of two sequential learning algorithms viz. Conditional random fields (CRF) and static vector machine (SVM) applied to the task of named entity recognition in Hindi. Since the features used are language independent hence the same procedure can be applied to tag the named entities for other Indian languages like Telgu, Bengali, Marathi etc. We have used CRF++ for implementing CRF algorithm and Yamcha for implementing SVM algorithm. The results show a superiority of CRF over SVM and are just a little lower than the highest results achieved for this task which is due to the non-usage of any pre-processing and post-processing steps. The system makes use of the contextual information of words along with various language independent features to label the named entities (NEs). We first present the two systems (CRF and SVM) and then compare their results for the same data.