Roman Urdu Sentiment Analysis Using Pre-trained DistilBERT and XLNet

Nikhar Azhar, Seemab Latif · 2022

Roman Urdu is a resource-poor language, therefore training a deep learning-based model from scratch is not that fruitful due to the lack of a large dataset. This is where the magic of transfer learning comes to the rescue. Using Huggingface's transformer models DistilBERT and XLNet we see a huge improvement in the results compared to the popular machine learning models Logistic Regression and Naïve Bayes. DistilBERT achieved 100% accuracy on our Roman Urdu dataset with only 2 epochs of fine-tuning while XLNet, though initially expected to perform better than the BERT model, achieved only about 86% accuracy with 4 epochs of fine-tuning.

Read the paper · More papers on PaperTik