Investigating Relationship between Data Augmentation Intensity and Model Performance in Natural Language Processing

Shouki Wada, Naoyuki Morimoto · 2024

In machine learning, data augmentation is known to contribute to improving model performance by increasing the quantity of data, particularly in cases where data collection is difficult. Widely utilized in image processing, it has also been studied in the context of natural language processing (NLP). The design of data augmentation strategies involves the optimization of techniques and the intensity of augmentation, etc. Therefore, investigating the relationship between data augmentation intensity and model performance is useful. In this study, following existing research, fidelity and diversity were utilized as indicators of data augmentation intensity in NLP, and the relationship between data augmentation intensity and model performance was investigated under various conditions. The results of the experiments showed cases where combining weak and strong augmentations led to improved model performance.

Read the paper · More papers on PaperTik