Deep Generative Data Augmentation for Natural Language Processing

유강민 · Seoul National University Open Repository (Seoul National University) · 2020

Recent advances in generation capability of deep learning models have spurred interest in utilizing deep generative models for unsupervised generative data augmentation (GDA).Generative data augmentation aims to improve the performance of a downstream machine learning model by augmenting the original dataset with samples generated from a deep latent variable model.This data augmentation approach is attractive to the natural language processing community, because (1) there is a shortage of text augmentation techniques that require little supervision and ( 2) resource scarcity being prevalent.In this dissertation, we explore the feasibility of exploiting deep latent variable models for data augmentation on three NLP tasks: sentence classification, spoken language understanding (SLU) and dialogue state tracking (DST), represent NLP tasks of various complexities and properties -SLU requires multi-task learning of text classification and sequence tagging, while DST requires the understanding of hierarchical and recurrent data structures.For each of the three tasks, we propose a task-specific latent variable model based on conditional, hierarchical and sequential variational autoencoders (VAE) for multi-modal joint modeling of linguistic features and the relevant annotations.We conduct extensive experiments to statistically justify our hypothesis that deep generative data augmentation is beneficial for all subject tasks.Our experiments show that deep generative data augmentation is effective for the select tasks, supporting the idea that the technique can potentially be utilized for other range of NLP tasks.Ablation and qualitative studies reveal deeper insight into the underlying mechanisms of generative data augmentation.As a secondary contribution, we also shed light onto the recurring posterior collapse phenomenon in autoregressive VAEs and, subsequently, propose novel techniques to reduce the model risk, i which is crucial for proper training of complex VAE models, enabling them to synthesize better samples for data augmentation.In summary, this work intends to demonstrate and analyze the effectiveness of unsupervised generative data augmentation in NLP.Ultimately, our approach enables standardized adoption of generative data augmentation, which can be applied orthogonally to existing regularization techniques.

Read the paper · More papers on PaperTik