Multilingual Named Entity Recognition in Tweets using Wikidata

Ramy Eskander, Peter Martigny, Shubhanshu Mishra · Zenodo (CERN European Organization for Nuclear Research) · 2020

This paper describes a simple way to improve performance of Named Entity Recognition systems across languages using knowledge from Wikidata on Social Media Corpus. We use dictionary based and zero-shot multilingual transfer. Abstract: We present Named Entity Recognition Models for Tweets trained without any gold data. Our method utilizes Wikidata as a gazetteer to generate a high precision dataset on which we train a BiLSTM CRF model in monolingual as well as multilingual setting. Our models when evaluated on multiple public benchmarks show that it is possible to generate high precision (but low recall) NER models for tweets across multiple languages using the entities in Wikidata. A video explaining our approach can be found on YouTube at: https://www.youtube.com/watch?v=GPFKnLAwmAc

Read the paper · More papers on PaperTik