XF2T: Cross-lingual Fact-to-Text Generation for Low-Resource Languages
Shivprasad Sagare, Tushar Abhishek, Bhavyajeet Singh, Anubhav Sharma, Manish Gupta, Vasudeva Varma · 2023
Multiple business scenarios require an automated generation of descriptive humanreadable text from structured input data.This has resulted into substantial work on fact-totext generation systems recently.Unfortunately, previous work on fact-to-text (F2T) generation has focused primarily on English mainly due to the high availability of relevant datasets.Only recently, the problem of crosslingual fact-to-text (XF2T) was proposed for generation across multiple languages alongwith a dataset, XALIGN for eight languages.However, there has been no rigorous work on the actual XF2T generation problem.We extend XALIGN dataset with annotated data for four more languages: Punjabi, Malayalam, Assamese and Oriya.We conduct an extensive study using popular Transformer-based text generation models on our extended multilingual dataset, which we call XALIGNV2.Further, we investigate the performance of different text generation strategies: multiple variations of pretraining, fact-aware embeddings and structure-aware input encoding.Our extensive experiments show that a multi-lingual mT5 model which uses fact-aware embeddings with structure-aware input encoding leads to best results (30.90 BLEU, 55.12 METEOR and 59.17 chrF++) across the twelve languages.We make our code and dataset publicly available 1 , and hope that this will help advance further research in this critical area. https://github.com/blitzprecision/ XAlignV22 A fact is a triple composed of subject, relation and object.