IndicBART: A Pre-trained Model for Indic Natural Language Generation
Raj Dabre, Himani Shrotriya, Anoop Kunchukuttan, Ratish Puduppully, Mitesh M. Khapra, Pratyush Kumar · Findings of the Association for Computational Linguistics: ACL 2022 · 2022
In this paper, we study pre-trained sequenceto-sequence models for a group of related languages, with a focus on Indic languages.We present IndicBART, a multilingual, sequenceto-sequence pre-trained model focusing on 11 Indic languages and English.IndicBART utilizes the orthographic similarity between Indic scripts to improve transfer learning between similar Indic languages.We evaluate In-dicBART on two NLG tasks: Neural Machine Translation (NMT) and extreme summarization.Our experiments on NMT and extreme summarization show that a model specific to related languages like IndicBART is competitive with large pre-trained models like mBART50 despite being significantly smaller.It also performs well on very low-resource translation scenarios where languages are not included in pre-training or fine-tuning.Script sharing, multilingual training, and better utilization of limited model capacity contribute to the good performance of the compact IndicBART model.