Improved Named Entity Tagset for Punjabi Language
Amandeep Kaur, Gurpreet Singh Josan · 2014
Annotated corpus plays an important role in developing machine learning based Named Entity Recognition system. For creating an annotated corpus, it is important to decide in advance the Named Entity Tagset to be used. A Named Entity Tagset is defined as a collection of tags or labels, in the form of a scheme, indicating the named entity class of a word to which it belongs in the text. In this paper we have proposed an improved Named Entity Tagset of 14 tags for the task of Named Entity Recognition in Punjabi Language. This improvement was realized from the challenges faced during annotation process in our previous research work with 12 tags. Apart from this we have discussed the importance and issues related to defining a Named Entity Tagset and annotation guidelines. We have also discussed various global tagsets found in Literature. We have referred Extended Named Entity Hierarchy for improving our current tagset.