Unveiling the Impact of Data Augmentations on Contemporary NLP Systems

Zhuofeng Wu · Deep Blue (University of Michigan) · 2024

Distilling human knowledge into machine learning models represents a formidable yet fascinating challenge within the realm of modern artificial intelligence. This potential paves the way for machines not just to mimic human language but to achieve a genuine comprehension of it. In this dissertation, we delve into the integration of human knowledge into textual representation models, particularly through the lens of data augmentations. On the one hand, given the heavy reliance of contemporary machine learning models on data, transforming human knowledge into a data-compatible format emerges as a pragmatic approach for today's systems. On the other hand, by augmenting existing datasets with human knowledge, we enable models to encounter and interpret a broader array of human-understood scenarios. This expansion of perspective significantly enhances the models' generalization capabilities, moving them closer to a true understanding of human language. This thesis offers a thorough examination of data augmentations within textual representation learning models. In its initial segment, we investigate data augmentations' impact on pre-trained language models, applying a range of NLP augmentations within a contrastive learning framework. This approach not only showcases a competitive edge but also bridges a vital gap in our comprehension of optimizing contrastive learning for language pre-training. In the thesis's second section, we introduce an innovative training methodology for contrastive learning, designed to enhance the application of data augmentation methods. This new approach is compatible with the latest contrastive learning models, providing a versatile framework for augmentations of all types. Further broadening the scope, the third part explores the versatility of data augmentation techniques beyond the confines of contrastive learning, demonstrating their effectiveness in various fine-tuning scenarios, both with and without labels. Initially, we introduce a novel augmentation concept known as the input-dependent prompt. Integrating this prompt with the original input significantly boosts model performance in environments where parameter modification is minimal, with most parameters fixed and only a few subject to tuning. Subsequently, we examine the efficacy of data augmentations in zero-shot scenarios during the test-time inference stage. These scenarios present unique challenges due to the absence of labels but also offer opportunities for the application of data augmentations. In summary, this thesis navigates through an extensive array of studies focusing on data augmentations across various stages of modern large language models, including pre-training, further pre-training, fine-tuning, and test time adaptation. Our exploration into embedding human knowledge within language models is pioneering, holding significant potential for advancing the domain of textual representation learning. We posit that our research direction not only breaks new ground but also illuminates promising pathways for the future evolution of language model development and application.

Read the paper · More papers on PaperTik