OntoHindi NER - An Ontology Based Novel Approach for Hindi Named Entity Recognition

Arti Jain, Devendra Kumar Tayal, Anuja Arora · 2018

Named Entity Recognition (NER) is defined as an identification and classification of Named Entities (NEs) in a given text. Now-a-days, language and domain-specific NER systems are in progress, therefore Ontology-based Hindi language NER methodology for health domain (OntoHindi NER) is proposed in this paper. Hindi Health Data (HHD) is crawled from Indian websites- Traditional Knowledge Digital Library, Ministry of Ayush, University of Patanjali, and Linguistic Data Consortium for Indian Languages. HHD corpus comprises 310,530 words having considered NEs as- Person (PER), Disease (DIS), Consumable (CNS) and Symptom (SMP). OntoHindi NER maps ontology for health data to recognize NEs by maintaining hierarchical information of the ontological category of HHD corpus words. OntoHindi NER comprises of six vital stages- preparing gazetteer lists (four initial seed lists that are extended using Hindi WordNet synset), HHD pre-processing, HHD feature engineering, string-matching based feature engineering (Levenshtein distance and Linguistic based Improved Lin’s Matcher (LILM)), Concept Hierarchy based Mapping (CHM), COncept Selection and Aggregation (COSA). CHM structuralizes ontological mapping (1:1, 1:m, and m:1) on the basis of ontological semantic relationship, and COSA formulates ontology-based NE clusters on HHD corpus through silhouette measure. Further, cluster aggregations are exploited using standard k-means clustering which clubs HHD corpus into varied NE clusters. OntoHindi NER performance evaluation is done using four standard measures- precision, recall, F-score and model fitting time. Results of NER for Hindi language are validated through 5-fold cross-validation.

Read the paper · More papers on PaperTik