Named Entity Recognition for Thai Historical Data
Nasith Laosen, Kanjana Laosen, Thummarat Paklao · 2024
Named Entity Recognition (NER) is a fundamental task in Natural Language Processing (NLP), enabling various advanced NLP applications like information extraction, question answering, and text summarization. Thai NER presents unique challenges due to the absence of capitalization, explicit word boundaries, and sentence-ending punctuation. While NER on general domain Thai datasets has been explored, its application to historical text remains an underexplored area. Historical texts contain specialized terminology, necessitating domain-specific NER solutions for optimal performance. This study investigates Thai historical NER, utilizing historical textual data collected from the Wikipedia website. Our goal is to identify suitable word segmentation and NER methods. We evaluate the performance of the Attacut word segmentation algorithm against Deepcut and Newmm. Our findings demonstrate Attacut's superiority, achieving an F1-score of 0.9557. Furthermore, our proposed SBC model (Sentence Transformer + BiLSTM + CRF) outperforms pre-trained LLMs (BERT-th1, XLM-R, WangchanBERTa) in NER, achieving an average F1-score of 0.97. The overall performance of the Attacut algorithm and the SBC model highlights their suitability for developing advanced NLP applications within the Thai historical domain.