Structure-Aware Joint POS and Morphological Tagging for Mongolian with Pretrained Models

Gang Jin, Na Ta, Qin Shu, Hasi, Celimuge Wu · 2025

As a typical agglutinative language, Mongolian presents significant challenges for joint part-of-speech (POS) and morphological tagging due to its complex morphological variations, flexible word order, and scarcity of annotated linguistic resources. This paper proposes SACMoTagger, a structure-aware joint tagging model specifically designed for Mongolian, which innovatively integrates a self-trained character-level pretrained language model (Mo-RoBERTa), bidirectional long short-term memory (BiLSTM), multi-head self-attention mechanisms, a Conditional Random Field (CRF)-based decoder, and introduces a label transition constraint mechanism to enhance the structural consistency of label sequences. Experimental results on a self-constructed fine-grained Mongolian POS and morphological tagging corpus demonstrate that the proposed model significantly outperforms traditional CRF-only, BiLSTM+CRF, and Transformer baseline models in terms of accuracy and F1-score. The proposed method provides an efficient and practical technical approach to structured tagging problems in low-resource languages.

Read the paper · More papers on PaperTik