IndoRobusta: Towards Robustness Against Diverse Code-Mixed Indonesian Local Languages
Muhammad Farid Adilazuarda, Samuel Cahyawijaya, Genta Indra Winata, Pascale Fung, Ayu Purwarianti · 2022
Significant progress has been made on Indonesian NLP.Nevertheless, exploration of the codemixing phenomenon in Indonesian is limited, despite many languages being frequently mixed with Indonesian in daily conversation.In this work, we explore code-mixing in Indonesian with four embedded languages, i.e., English, Sundanese, Javanese, and Malay; and introduce IndoRobusta 1 , a framework to evaluate and improve the code-mixing robustness.Our analysis shows that the pre-training corpus bias affects the model's ability to better handle Indonesian-English code-mixing when compared to other local languages, despite having higher language diversity.