IndoAbbr: A New Benchmark Dataset for Indonesian Abbreviation Identification
Shanshan Li, Nankai Lin, Lixian Xiao, Shengyi Jiang · 2020
Automatic abbreviations identification is an essential task for natural language processing on Indonesian, such as named entity recognition, information retrieval and question answering. One of the challenges is the publicly available and large-scale dataset that is relatively rare and difficult to construct. The problem is even worse for low-resource languages such as Indonesian. In order to bridge the gap of Indonesian abbreviations identification, we constructed a dataset for Indonesian abbreviation identification and compared different methods on Indonesian abbreviation identification. Besides, we extracted three external features, respectively capitalization, consonant and whether appears in the Indonesian dictionary to improve the performance of the models. The results demonstrate that the external features while using the Bi-LSTM-CRF model and CNN-Bi-LSTM-CRF model are effective. Base on the dataset for Indonesian abbreviation identification mentioned above, the BERT model shows the best performance, achieving the average accuracy of 98.28% on the dataset.