RELATION EXTRACTION DATASET FOR THE RUSSIAN
RANEPA, Moscow, Russia, Denis Gordeev, Adis Davletov, RANEPA, Moscow, Russia, Alexey Rey, RANEPA, Moscow, Russia, G. R. Akzhigitova, RANEPA, Moscow, Russia, G. A. Geymbukh, RANEPA, Moscow, Russia · Computational Linguistics and Intellectual Technologies · 2020
There are few existing relation extraction datasets for the Russian language and they contain a rather small number of examples. Thus, we decided to create a new Ontonotes-based named entities and relation extraction sentence-level dataset called RURED. The dataset contains more than 500 annotated texts and more than 5,000 labelled relations. We also publish baseline models for relation extraction and named entity recognition trained on the dataset. Our models achieve 0.85 for named entity recognition and 0.78 for relation extraction in F1-score.