Two-stage approach in Russian named entity recognition

V. A. Mozharova, Natalia Loukachevitch · 2016

In this article we consider a two-stage prediction approach for named entity recognition in Russian. In the first stage, named entities are extracted by a machine learning method. After that our system collects the statistics of token classes and transforms this statistics to a feature set, which is used for training a new classifier. We consider three types of the two-stage features: the previous history, the whole document statistics, and global statistics of the whole collection. We carry out our experiments on several text collections. We show that the utilizing of the two-stage prediction approach improves the quality of named entity recognition.

Read the paper · More papers on PaperTik