A Bilingual Legal NER Dataset and Semantics-Aware Cross-Lingual Label Transfer Method for Low-Resource Languages

Paerhati Tulajiang, Yuanyuan Sun, Yuanyu Zhang, Yingying Le, Kelaiti Xiao, Hongfei Lin · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025

Named Entity Recognition (NER) in specialized domains for low-resource languages remains a significant challenge due to data scarcity and the complexity of domain-specific terminology. Existing cross-lingual approaches—spanning model-transfer and data-transfer paradigms—often suffer from semantic drift and inadequate domain adaptation. To address these limitations, we introduce BiLegalNERD, the first bilingual Chinese–Uyghur legal NER dataset, constructed via a semantics-aware label transfer strategy. We further propose CUTLM, a cross-lingual annotation method that combines dual translation with Levenshtein-based alignment to ensure high-fidelity preservation of entity boundaries across languages. In addition, we present BiLegalNER, a domain-adapted multilingual NER model incorporating vocabulary expansion and bilingual fine-tuning, significantly enhancing performance on Uyghur legal texts. Experiments demonstrate that BiLegalNER achieves state-of-the-art results, with F1-scores of 86.65% on automatically generated training data and 89.11% on fully human-annotated data—outperforming the strongest multilingual baseline by 4.18% and 4.64%, respectively. Moreover, CUTLM surpasses prior cross-lingual transfer methods by up to 9.89%, confirming its effectiveness in preserving entity integrity during label projection. These findings establish a new benchmark for Uyghur legal NER and provide a scalable framework for cross-lingual NER in low-resource settings.

Read the paper · More papers on PaperTik