Research on improving the generalization ability of NLP categorization models: A comparison of Mixup and generative adversarial networks
Jiyan Chen · Applied and Computational Engineering · 2023
The application of deep neural networks in tasks related to natural language processing has grown in recent years. However, training deep learning models often requires a large amount of data to produce accurate and generalizable results. When the data are insufficient, categorization models may suffer from overfitting or poor performance. This paper compares two methods for enhancing the generalization ability of natural language processing categorization models when data are insufficient: Mixup and Generative Adversarial Networks (GANs). This paper uses the Internet Movie Database (IMDB) dataset for a binary classification task, Mixup method, and Relational Generative Adversarial Networks (RelGAN) model to generate new data, respectively, and compares them with the original dataset. Experimental results indicate that Mixup method can reduce overfitting risk effectively and enhance the model’s robustness, while the GAN model does not show obvious advantages. This paper supports the idea that using noise-adding operations to enhance the dataset is feasible in the case of small samples rather than using generative models.