QazNLP: Constraint-Aware Multi-Task Sequence Labeling for Morphologically Rich Low-Resource Languages

Айгерим Айтим · IEEE Access · 2026

Automatic processing of morphologically rich, agglutinative, and low-resource languages remains challenging because productive affixation increases lexical sparsity, weakens statistical generalization, and often produces inconsistent predictions across related linguistic annotation tasks. This study presents QazNLP, a constraint-aware multi-task framework for Kazakh that jointly performs morphological tagging, part-of-speech tagging, and named entity recognition using a shared transformer encoder with task-specific prediction heads. To improve structural reliability, the framework introduces differentiable cross-task compatibility penalties that discourage linguistically invalid label combinations during training and constrained decoding. The study further provides a reproducible evaluation setting based on a cleaned Kazakh news corpus with fixed data splits and robustness diagnostics for out-of-vocabulary tokens, long agglutinative word forms, reduced-data regimes, and cross-task contradiction analysis. In addition, the manuscript explicitly documents the encoder and tokenization setup, justifies the choice of pretrained encoder, reports decoding complexity, and situates the NER component with respect to the KazNERD benchmark. Experimental results show that the proposed model consistently outperforms competitive single-task and shared multi-task baselines in joint average F1, robustness under sparse-data conditions, and structural consistency of predicted labels. The findings indicate that explicit compatibility-aware optimization offers a practical and extensible direction for sequence labeling in morphologically rich low-resource languages.

Read the paper · More papers on PaperTik