RELATION EXTRACTION DATASET FOR THE RUSSIAN

RANEPA, Moscow, Russia, Denis Gordeev, Adis Davletov, RANEPA, Moscow, Russia, Alexey Rey, RANEPA, Moscow, Russia, G. R. Akzhigitova, RANEPA, Moscow, Russia, G. A. Geymbukh, RANEPA, Moscow, Russia · Computational Linguistics and Intellectual Technologies · 2020

There are few existing relation extraction datasets for the Russian language and they contain a rather small number of examples. Thus, we decided to create a new Ontonotes-based named entities and relation extraction sentence-level dataset called RURED. The dataset contains more than 500 annotated texts and more than 5,000 labelled relations. We also publish baseline models for relation extraction and named entity recognition trained on the dataset. Our models achieve 0.85 for named entity recognition and 0.78 for relation extraction in F1-score.

Read the paper · More papers on PaperTik