Named Entity Recognition in Bengali Text Using Merged Hidden Markov Model and Rule Base Approach
Mah Dian Drovo, Moithri Chowdhury, Saiful Islam Uday, Amit Kumar Das · 2019
Named Entity Recognition (NER) is the subtask of Natural Language Processing (NLP) which tries to achieve human level on a specific domain (e.g. newspaper) to identify named entities. It seeks to locate and classify named entities (Person Name, Location, Organization names etc.), which is the most vital step of Information Extraction (IE). In many cases Machine Learning (ML) is mostly used to perform NER. Apart from that, another method is applied which is known as Rule Base approach. This paper presents a method which is using both ML and Rule Base approach together for NER basing on Bengali language. Mainly the rule based approach has been merged with ML. For ML Hidden Markov Model (HMM) and for rule base approach Regular Expression has been used. A Named Entity (NE) tagged corpus has been developed by using Bengali newspaper, which consists of 10k words that has been manually annotated with seven tags. This paper concludes with experimental results which shows two distinctive ways of our proposed model.