Multilingual Named Entity Recognition Model for Location and Time Extraction of Forest Fire

Hafidz Sanjaya, Kusrini Kusrini, Kumara Ari Yuana, José Ramón Martínez Salio · 2024

The scarcity and limited access to multilingual datasets can hinder the training of named entity recognition models, especially for the application of social media sensing in disaster management, such as forest fires. Additionally, creating high-quality datasets also requires significant resources and efforts. Overcoming the limited availability of multilingual named entity recognition datasets in disaster management is a possible approach to reducing resources and efforts in order to make it more efficient. Therefore, this paper will compare the performance of several pre-trained multilingual BERT-based models on a web-based general-purpose dataset in Indonesian to extract the location and time of forest fires from social media text like Twitter. The fine-tuning results show that XLMRoBERTa (XLM-R) obtains the best performance with a validation loss of 0.0567, precision of 0.92, recall of 0.93, F1score of 0.92 and accuracy of 0.98. Testing results also show that XLM-R achieves the best overall performance on each label entity. The tweet validation results also obtained an acceptable accuracy of more than $50 \%$ for languages that were not in the main dataset.

Read the paper · More papers on PaperTik