Detecting Entailment in Code-Mixed Hindi-English Conversations
Sharanya Chakravarthy, Anjana Umapathy, Alan W. Black · 2020
The presence of large-scale corpora for Natural Language Inference (NLI) has spurred deep learning research in this area, though much of this research has focused solely on monolingual data.Code-mixing is the intertwined usage of multiple languages, and is commonly seen in informal conversations among polyglots.Given the rising importance of dialogue agents, it is imperative that they understand code-mixing, but the scarcity of code-mixed Natural Language Understanding (NLU) datasets has precluded research in this area.The dataset by Khanuja et al. (2020a) for detecting conversational entailment in codemixed Hindi-English text is the first of its kind.We investigate the effectiveness of language modeling, data augmentation, translation, and architectural approaches to address the codemixed, conversational, and low-resource aspects of this dataset.We obtain +8.09% test set accuracy over the current state of the art.