Exploring Deep Learning Methods for Text Augmentation to Handle Imbalanced Datasets in Natural Language Processing
Reeba Khan, Anoushka Venugopal · 2024
Exploring emotional nuances like joy, sadness, anger, and surprise in text data poses one of the most challenging problems in Natural Language Processing. Coupled with limited annotated datasets and an imbalance in datasets, it sounds like a recipe for machine learning disaster. Both the quantity and quality of the dataset is a concerned factor when building machine learning models. But there's always a challenge in finding sufficient data to train and test. This brings forth the problem of imbalanced datasets. This study addresses the challenge of imbalanced datasets in Natural Language Processing by exploring deep-learning-based text augmentation methods. The primary aim is to enhance the performance of NLP models by increasing the quantity and diversity of training data without compromising the original semantic content. Various techniques like BERT, Conditional Variational Autoencoders, Generative Adversarial Networks, Long Short-Term Memory Networks, and Word Embeddings, were evaluated for their efficacy in generating synthetic training data while preserving the semantics. The methodology involved a comparative analysis based on factors including data quality, training time, and robustness. Our results indicate that even though each model displays unique strengths and limitations, BERT outperform others in generating high-quality text data while preserving context and semantics. The findings suggest that deep learning with text augmentation presents a potent solution to address class imbalance.