NLP_DI at NADI 2024 shared task: Multi-label Arabic Dialect Classifications with an Unsupervised Cross-Encoder

Vani Kanjirangat, Tanja Samardżić, Ljiljana Dolamic, Fabio Rinaldi · 2024

We report the approaches submitted to the NADI 2024 Subtask 1: Multi-label countrylevel Dialect Identification (MLDID).The core part was to adapt the information from multiclass data for a multi-label dialect classification task.We experimented with supervised and unsupervised strategies to tackle the task in this challenging setting.Under the supervised setup, we used the model trained using NADI 2023 data and devised approaches to convert the multi-class predictions to multi-label by using information from the confusion matrix or calibrated probabilities.Under unsupervised settings, we used the Arabic-based sentence encoders and multilingual cross-encoders to retrieve similar samples from the training set, considering each test input as a query.The associated labels are then assigned to the input query.We also tried variations, such as cooccurring dialects derived from the provided development set.We obtained the best validation performance of 48.5% F-score using one of the variations with an unsupervised approach and the same approach yielded the best test result of 43.27% (Ranked 2).

Read the paper · More papers on PaperTik