Datasets and Performance Metrics for Greek Named Entity Recognition
Nikos Bartziokas, Thanassis Mavropoulos, Constantine L. Kotropoulos · 2020
The quality of linguistic resources determines the robustness and efficiency of language processing models to a great extent. This is especially evident with less represented languages, where dataset availability is scarce, preventing the development of state-of-the-art natural language processing algorithms. In this paper, we attempt to partially remedy the situation for the Greek language by introducing a dataset focused on named entity recognition (NER), coined as elNER. Motivated by CoNLL-2003 and OntoNotes 5 datasets in English, two versions of elNER are described, namely elNER-4 and elNER-18, respectively. Current state-of-the-art NER models are trained on these datasets. The performance metrics disclosed here are comparable to those achieved by the respective approaches in English.