Enriching the E2E dataset

Thiago Castro Ferreira, Helena Vaz, Brian Davis, Adriana Silvina Pagano · 2021

This study introduces an enriched version of the E2E dataset, one of the most popular language resources for data-to-text NLG.We extract intermediate representations for popular pipeline tasks such as discourse ordering, text structuring, lexicalization and referring expression generation, enabling researchers to rapidly develop and evaluate their data-totext pipeline systems.The intermediate representations are extracted by aligning nonlinguistic and text representations through a process called delexicalization, which consists in replacing input referring expressions to entities/attributes with placeholders.The enriched dataset is publicly available.1

Read the paper · More papers on PaperTik