Synonym-based Text Generation in Restructuring Imbalanced Dataset for Deep Learning Models
Febi Siti Sutria Ningsih, Purnomo Husnul Khotimah, Andria Arisal, Andri Fachrur Rozie, Devi Munandar, Dianadewi Riswantini, Ekasari Nugraheni, Wiwin Suwarningsih, Dian Kurniasari · 2022
One of which machine learning data processing problems is imbalanced classes. Imbalanced classes could potentially cause bias towards the majority classes due to the nature of machine learning algorithms that presume that the object cardinality in classes is around similar number. Oversampling or generating new objects in minority class are common approaches for balancing the dataset. In text oversampling method, semantic meaning loses often occur when deep learning algorithms are used. We propose synonym-based text generation for restructuring the imbalanced COVID-19 online-news dataset. Three deep learning models (MLP, CNN, and LSTM) using TF/IDF and word embedding (WE) feature are tested with the original and balanced dataset. The results indicate that the balance condition of the dataset and the use of text representative features affect the performance of the deep learning model. Using balanced data and deep learning models with WE greatly affect the classification significantly higher performances as high as 4%, 5%, and 6% in accuracy, precision, recall, and f1-score, respectively.