FEATURES FOR NAMED ENTITY RECOGNITION IN CZECH LANGUAGE
Pavel Král · 2011
This paper deals with Named Entity Recognition (NER). Our work focuses on the application for the Czech News Agency (CTK). We propose and implement a Czech NER system that facilitates the data searching from the CTK text news databases. The choice of the feature set is crucial for the NER task. The main contribution of this work is thus to propose and evaluate some different features for the named entity recognition and to create an “optimal” set of features. We use Conditional Random Fields (CRFs) as a classifier. Our system is tested on a Czech NER corpus with nine main named entity classes. We reached 58% of the F-measure with the best feature set which is sufficient for our target application.