Comparison of Pretrained BERT Embedding and NLTK Approach for Easy Data Augmentation in Research Title Document Classification Task
Muhammad Fikri Hasani, Kenny Jingga, Kristien Margi Suryaningrum, Michael Caleb · 2022
Data imbalance has been a problem in classification tasks, be it a binary classification or multiclass classification. Easy data augmentation is an approach to do data augmentation in natural language processing data by substituting term with its synonym, inserting a random synonym, or deletion. This research aims to see the performance of machine learning model if term inserted or changed is using BERT rather than NLTK. This research also proposed the hybrid approach of BERT and NLTK, and proposed sentence selection methods such as whole data and random pick. This research showed that EDA using only BERT gives the best performance for all sentence selection approach. However, several classes still showed reduced performance.