Sequence labelling for code-mixed text by an averaged perceptron approach
Daksh Sawhney · Oxford Journal of Student Scholarship · 2026
In mixed-language settings, NLP models face challenges because of intertwined grammatical structures and language discrepancies.This paper analyses a Part-of-Speech tagger used on Hindi-English code-mixed (Hinglish) data and compares it with monolingual English and Hindi data.The model achieved an accuracy of 93.60% on the English dataset, accuracy of 95.95% on the Hindi dataset, and 94.89% on the Hinglish dataset, respectively.This showed strong model performance in both monolingual and code-mixed settings.The study also tried to understand the patterns in sentence structure that led to tagging errors.It was found that the most misclassifications occurred between nouns, verbs, proper nouns, and adjectives because of their overlapping lexical forms.The analysis of sentence length indicated that medium-length sentences yielded marginally higher accuracy than long sentences.In contrast, the error distribution per sentence showed that most sentences contain zero or one tagging error, with a cascading effect observed in a minority of cases where a single error in a sentence led to more errors and isolated errors.In the future, researchers can focus on understanding the reason for dual language error patterns, such as common misclassifications in a complex sentence.