A Proposed Model for Bengali Named Entity Recognition Using Maximum Entropy Markov Model Incorporated with Rich Linguistic Feature Set
Fahmida Alam, Md Asiful Islam · 2020
Named Entity Recognition (NER) is a subpart of Information Extraction task that helps to find named entities and their class in a document. This is one of the fundamental tasks for many natural language processing tasks. NER categorizes the named entities in some predefined class like person, name, organization, date, number etc. In this paper, we have proposed a model to recognize the named entities from Bengali language. Three major approaches are followed to recognize named entities. They are rule-based approach, machine learning based approach and hybrid approach. Most of the existing Bengali named entity recognition model followed machine learning approaches. Existing work done on Conditional Random Forest (CRF), Support Vector Machine (SVM) and Hidden Markov Model (HMM) etc. As. Bengali is a complex language enriched with complex linguistic feature set., in this paper we have adopted a Maximum Entropy Markov Model (MEMM) based machine learning approach that can deal with complex sequences better than other approaches to classify the named entities from Bengali language. Also, we have described a rich linguistic feature set for training our model.