A new dataset for sentence-level complexity in Russian

Vladimir Vladimirovich Ivanov, Elbayoumi Mohamed Gamal · Computational Linguistics and Intellectual Technologies · 2023

Text complexity prediction is a well-studied task. Predicting complexity sentence-level has attracted less research interest in Russian. One possible application of sentence-level complexity prediction is more precise and fine-grained modeling of text complexity. In the paper we present a novel dataset with sentence-level annotation of complexity. The dataset is open and contains 1,200 Russian sentences extracted from SynTagRus treebank. Annotations were collected via Yandex Toloka platform using 7-point scale. The paper presents various linguistic features that can contribute to sentence complexity as well as a baseline linear model.

Read the paper · More papers on PaperTik