Large Language Model Data Augmentation for Text-Pair Classification Tasks
Yuyang Li, Yuqing Zhang, Zelin Du, Ziqi Guo · 2024
In recent years, large language models (LLMs) have demonstrated remarkable capabilities across various natural language processing tasks.This study explores the application of LLMs for data augmentation in text-pair classification tasks, such as semantic textual similarity and natural language inference.We propose a novel framework that leverages LLMs to generate diverse and contextually relevant paraphrases and transformations of text-pairs, enhancing the training data without manual annotation effort.Our experiments on widely-used benchmarks show that the augmented data not only improves model performance but also increases its robustness to out-of-domain examples.We perform extensive ablation studies to understand the contribution of different augmentation strategies and analyze the trade-offs between data diversity and noise.Additionally, we assess the generalization capabilities of models trained with augmented data across multiple architectures and dataset sizes.The results suggest that LLM-driven data augmentation is a promising approach to overcome data scarcity, reduce overfitting, and enhance the adaptability of text-pair classification systems.