Named Entity Recognition for South and South East Asian Languages: Taking Stock

Anil Kumar Singh · 2008

In this paper we first present a brief discussion of the problem of Named Entity Recognition (NER) in the context of the IJCNLP workshop on NER for South and South East Asian (SSEA) languages1. We also present a short report on the development of a named entity annotated corpus in five South Asian language, namely Hindi, Bengali, Telugu, Oriya and Urdu. We present some details about a new named entity tagset used for this corpus and describe the annotation guidelines. Since the corpus was used for a shared task, we also explain the evaluation measures used for the task. We then present the results of our experiments on a baseline which uses a maximum entropy based approach. Finally, we give an overview of the papers to be presented at the workshop, including those from the shared task track. We discuss the results obtained by teams participating in the task and compare their results with the baseline results. 1

Read the paper · More papers on PaperTik