Research on Data Augmentation Techniques for Text Classification Based on Antonym Replacement and Random Swapping

Shaoyan Wang, Yu Xiang · 2024

Traditional simple data augmentation techniques have been proven to be effective in enhancing the performance of models. Among these techniques, researchers have explored the use of synonym replacement and random position swapping. However, there has been limited exploration of the augmentation technique involving antonym replacement, and the focus of random swapping has mostly been on the swapping of positions between two words. Additionally, an important challenge in data augmentation is determining which data should be augmented. In this paper, we propose two data augmentation techniques: antonym replacement for data at a moderate difficulty level and random position swapping based on specific positions and proportions. We investigate the impact of these augmentation techniques on the performance of text classification models. Specifically, for the augmented samples obtained through antonym replacement, we propose using similarity and predictive models to assign labels. For random position swapping, we primarily explore the swapping of word positions within sentences and different swapping methods. Through these two augmentation techniques, we expand our limited text data and achieve improved performance on classification tasks.

Read the paper · More papers on PaperTik