IndicXNLI: Evaluating Multilingual Inference for Indian Languages
Divyanshu Aggarwal, Vivek Gupta, Anoop Kunchukuttan · 2022
While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited.To this end, we introduce INDICXNLI, an NLI dataset for 11 Indic languages.It has been created by high-quality machine translation of the original English XNLI dataset and our analysis attests to the quality of INDICXNLI.By finetuning different pre-trained LMs on this IN-DICXNLI, we analyze various cross-lingual transfer techniques with respect to the impact of the choice of language models, languages, multi-linguality, mix-language input, etc.These experiments provide us with useful insights into the behaviour of pre-trained models for a diverse set of languages.