IndoBERT Based Data Augmentation for Indonesian Text Classification
Fuad Muftie, Muhammad Haris · 2023
Text data augmentation has been able to improve the performance of models or algorithms for text classification and sentiment analysis. In major cases, rule-based augmentation techniques can be easily applied to various languages such as Indonesia language. However, language model-based augmentation techniques rarely investigated for Indonesian language data. Therefore, this paper presents a text data processing model for Indonesian language, which has limited data, to perform text preprocessing and data augmentation techniques by selectively inserting words based on IndoBERT. This IndoBERT-based augmentation is able to generate data that still retains meaning and sentiment similar to the original data. The testing of this Twitter text dataset yielded results showing that the proposed augmentation technique was able to increase accuracy and outperform the Random Insert augmentation technique.