A Data Bootstrapping Recipe for Low-Resource Multilingual Relation Classification
Arijit Nag, Bidisha Samanta, Animesh Mukherjee, Niloy Ganguly, Soumen Chakrabarti · 2021
Relation classification (sometimes called 'ex traction') requires trustworthy datasets for fine tuning large language models, as well as for evaluation.Data collection is challenging for Indian languages, because they are syntacti cally and morphologically diverse, as well as different from resourcerich languages like En glish.Despite recent interest in deep gen erative models for Indian languages, relation classification is still not wellserved by pub lic data sets.In response, we present IndoRE, a dataset with 21K entity and relationtagged gold sentences in three Indian languages, plus English.We start with a multilingual BERT (mBERT) based system that captures entity span positions and type information, and pro vides competitive monolingual relation clas sification.Using this system, we explore and compare transfer mechanisms between lan guages.In particular, we study the accuracy efficiency tradeoff between expensive gold in stances vs. translated and aligned 'silver' in stances.We release the dataset for future re search. 1