Fine-tuning Urdu NER Models Using Context-Aware Embeddings

Nikhar Azhar, Seemab Latif, Sahar Arshad · 2024

Many applications rely on Named Entity Recognition (NER) to accurately identify and extract key pieces of information from large amounts of text. For applications that deal with high-risk and sensitive data, having a reliable and efficient model becomes crucial. While there have been many advancements for NER in the English language, low-resource languages like Urdu are still lagging behind. In this research, we have focused on Urdu text NER. We used the MK-PUCIT corpus, which is currently the largest labeled dataset available for Urdu NER. We fine-tuned two transformer models, Urduhack's “roberta-urdu-small” and Google's “muril-based-cased”, and achieved Fl-scores of 0.942 and 0.941 respectively. However, through analyzing the results of our fine-tuning, we found many mislabeling errors in the MK-PUCIT corpus. These errors include incorrect tagging of names as locations and punctuation marks as entities. Despite these issues, our study shows the potential for enhancing Urdu NER by addressing these errors and improving the quality of the dataset.

Read the paper · More papers on PaperTik