Few-Shot and One-Shot POS Tagging for Dogri: A Data-Efficient Approach

Vipul Saluja, Jyotshna Dongardive · Procedia Computer Science · 2026

Part-of-Speech (POS) tagging plays an important role in Natural Language Processing (NLP). But its progress for low-resource and morphologically complex languages such as Dogri remains slow. The low availability of annotated data limits the effectiveness of conventional machine learning models. This study explores whether reliable POS tagging can be achieved with low training data through one-shot and few-shot learning approaches. We investigate meta-learning models (Prototypical Networks and MAML), multilingual transformer architectures (mBERT, IndicBERT), and cross-lingual transfer from related Indic languages. The performance over zero-shot settings have been improved significantly with limited number of examples while annotating Dogri corpus with the ILPOSTS tagset. IndicBERT attains close to full-supervision performance with just 20 annotated samples per tag, achieving around 84% accuracy and 81% Macro-F1. Analysis across tags has shown strong learning for open-class words such as nouns, verbs, and pronouns as compared to auxiliaries and postpositions. These findings confirm that practical Dogri POS taggers can be built with minimal data. Thus offering a scalable model for other low-resource Indic languages.

Read the paper · More papers on PaperTik