IrekiaLFes: a New Open Benchmark and Baseline Systems for Spanish Automatic Text Simplification

Itziar González-Dios, Iker Gutiérrez-Fandiño, Oscar M. Cumbicus-Pineda, Aitor Soroa · 2022

Automatic Text simplification (ATS) seeks to reduce the complexity of a text for a general or a target audience.In the last years, deep learning methods have become the most used systems in ATS research, but these systems need large and good-quality datasets to be evaluated.Moreover, these data are available on a large scale only for English and in some cases with restrictive licenses.In this paper, we present IrekiaLF_es, an open-license benchmark for Spanish text simplification.It consists of a document-level corpus and a sentence-level test set that has been manually aligned.We also conduct a neurolinguistically-based evaluation of the corpus in order to reveal its suitability for text simplification.This evaluation follows the Lexicon-Unification-Linearity (LeULi) model of neurolinguistic complexity assessment.Finally, we present a set of experiments and baselines of ATS systems in a zero-shot scenario.

Read the paper · More papers on PaperTik