Projecting named entity recognizers without annotated or parallel corpora

Jue Hou, Maximilian Koppatz, José María Hoya Quecedo, Roman Yangarber · DSpace repository (University of Tartu) · 2019

Named entity recognition (NER) is a task extensively researched in the field of NLP.NER typically requires large annotated corpora for training usable models.This is a problem for languages which lack large annotated corpora, such as Finnish.We propose an approach to create a named entity recognizer for Finnish by leveraging preexisting strong NER models for English, with no manually annotated data and no parallel corpora.We automatically gather a large amount of chronologically matched data in the two languages, then project named entity annotations from the English documents onto the Finnish ones, by resolving the matches with simple linguistic rules.We use this "artificially" annotated data to train a BiLSTM-CRF NER model for Finnish.Our results show that this method can produce annotated instances with high precision, and the resulting model achieves state-of-the-art performance.

Read the paper · More papers on PaperTik