Studying the impact of various features on the performance of Conditional Random Field-based Arabic Named Entity Recognition

Alia Morsi, Ahmed Rafea · 2013

The task of Named Entity Recognition (NER) is crucial to Natural Language Processing (NLP). NER can be defined as the computational identification and classification of Named Entities in running text. The importance of NER stems from the variety of Natural Language Processing applications where accurate NLP would be highly useful. Such include machine translation and information extraction. In this paper, we present a series of experiments to explore the impact of using various feature sets on NER results for Modern Standard Arabic (MSA) text. We rely on language independent features and we employ CRF based models for all our experiments. We create our own baseline model to use its results for comparison. Our best result is a 68.05 F-Measure, which is 10.82 points above our baseline.

Read the paper · More papers on PaperTik