Error Analysis for POS Tagging of Hindi-English Code-Mixed Data
Aashi Tiwari, Xiaoxi Luo · 2025
The phenomenon of code mixing (CM), particularly between Hindi and English, is increasingly prevalent in digital communication, especially on social media. The linguistic features of CM content often contain informal grammar, spelling errors, transliteration, etc., which makes the task of Natural Language Processing (NLP) of such content hard. This paper discusses the implementation of a Part-of-Speech (POS) tagger built using a single-layer Perceptron for CM content on social media and analyses the sources of inaccuracies in tagging. The tagger achieves an accuracy of 84%. The study tries to understand the types of sentence patterns that lead to tagging errors. It shows the correlation of tagging errors with neighbouring word context, sentence length, etc. Certain misclassifications are found to be more common than others, for example, nouns for verbs, which may be related to the difference in the way the two languages place verbs in sentences. POS tagging errors have also been shown to have a cascading effect: one error in a sentence leads to more, and isolated errors are less common. Understanding these error patterns, including common misclassifications, shows how these errors are linked to linguistic features like sentence structure, and highlights areas of improvement for further research.