Initial Explorations on using CRFs for Turkish Named Entity Recognition
Gökhan Akın Åžeker, GulÅŸen EryiÄŸitIstanbul Technical University · International Conference on Computational Linguistics · 2012
This paper reports the highest results (95% in MUC and 92% in CoNLL metric) in the literature for Turkish named entity recognition; more specifically for the task of detecting person, location and organization entities in general news texts. We give an in depth analysis of the previous reported results and make comparisons with them whenever possible. We use conditional random fields (CRFs) as our statistical model. The paper presents initial explorations on the usage of rich morphological structure of the Turkish language as features to CRFs together with the use of some basic and generative gazetteers.