Classification of URDU headline news using Bidirectional Encoder Representation from Transformer and Traditional Machine learning Algorithm

Kainat Mujahid, Sania Bhatti, Mohsin Ali Memon · 2021

the web is full of unstructured data and getting the right information from that is a crucial problem. Recent advancements in the field of data mining on English Corpus provided the opportunity to automatically classify the text from different sources into broad categories. However, due to the lack of Urdu language resources, less attention has been paid to classifying text written in Urdu. Urdu is one of the most commonly spoken languages in Asian countries and also a national language of Pakistan. Urdu is a morphologically rich language, contains special diacritics, and free word order which makes it a difficult task for automatic text classification. In this work, we introduce the fine-tuning BERT architecture for the classification of Urdu headline news. Also, perform Logistics regression (LR) and multilayer perceptron (MLP) with Bert features. The proposed model classifies the Urdu headline news based on their predefined labels. We performed extensive experiments on Urdu news dataset iNLTK. Our proposed BERT-fine-tuned model outperforms baseline MLP and LG models and produces average accuracy for BERT, MLP, and LG with 95%, 94%, and 93% respectively.

Read the paper · More papers on PaperTik