A dataset of Tibetan named entity recognition—TibNER

Mengting Zhou, Cairang EJIAN, Caidan DAOJI, Xiaoke Qi, Xiaobing ZHAO · China Scientific Data · 2024

Structured linguistic resources are an important foundation for natural language processing. Due to the lack of publicly available large-scale datasets, few researches focus on Tibetan Named Entity Recognition, with limited achievements. Therefore, in this paper we semi-automatically construct a dataset of Tibetan Named Entity Recognition named “TibNER” based on an entity dictionary. To ensure the quality of the dataset, we manually proofread the automatic annotation results. TibNER contains 20,096 sentences with an average sentence length of 44.2069 syllables, and the labelled entities include speaker names, place names and organization names, with a total of 43,678 entities across these categories. In order to validate the dataset, we tested it on three mainstream sequence annotation models, and the best model had an F1 value of 80.60%. According to studies, this dataset can not only provide data construction experience for low-resource languages, but also serves as a solid data foundation for researches like Tibetan named entity recognition.

Read the paper · More papers on PaperTik