Peer Review #1 of "Multi-label emotion classification of Urdu tweets (v0.1)"

2022

Urdu is a widely used language in South-Asia and worldwide.While there are similar datasets available in English, we created the first multi-label emotion dataset consisting of 6,043 tweets and six basic emotions in the Urdu Nastalíq script.A Multi-Label (ML) classification approach was adopted to detect emotions from Urdu.The morphological and syntactic structure of Urdu makes it a challenging problem for multi-label emotion detection.In this paper, we build a set of baseline classifiers such as machine learning algorithms (Random forest (RF), Decision tree (J48), Sequential minimal optimization (SMO), AdaBoostM1, and Bagging), deep-learning algorithms (Convolutional Neural Networks (1D-CNN), Long short-term memory (LSTM), and LSTM with CNN features) and transformer-based baseline (BERT).We used a combination of text representations: stylometric-based features, pre-trained word embedding, word-based n-grams, and character-based n-grams.The paper highlights the annotation guidelines, dataset characteristics and insights into different methodologies used for Urdu based emotion classification.We present our best results using Micro-averaged F1, Macro-averaged F1, Accuracy, Hamming Loss (HL) and Exact Match (EM) for all tested methods.

Read the paper · More papers on PaperTik