Meta-Learning for Offensive Language Detection in Code-Mixed Texts

Gautham Vadakkekara Suresh, Bharathi Raja Chakravarthi, John Philip McCrae · Forum for Information Retrieval Evaluation · 2021

This research investigates the application of Model-Agnostic Meta-Learning (MAML) and ProtoMAML to identify offensive code-mixed text content on social media in Tamil-English and Malayalam-English code-mixed texts. We follow a two-step strategy: The XLM-RoBERTa (XLM-R) model is trained using the meta-learning algorithms on a variety of tasks having code-mixed data, monolingual data in the same language as the target language and related tasks in other languages. The model is then fine-tuned on target tasks to identify offensive language in Malayalam-English and Tamil-English code-mixed texts. Our results show that meta-learning improves the performance of models significantly in low-resource (few-shot learning) tasks1. We also introduce a weighted data sampling approach which helps the model converge better in the meta-training phase compared to traditional methods.

Read the paper · More papers on PaperTik