On the Effectiveness of Graph Data Augmentation for Source Code Learning

Zeming Dong, Qiang Hu, Zhenya Zhang, Jianjun Zhao · 2023

The methodology that uses deep learning to solve software engineering tasks, such as bug detection, is known as source code learning. Due to the graph nature of source code, graph learning, empowered by graph neural networks (GNNs), has been increasingly adopted for source code learning. Like other deep learning contexts, source code learning also relies on massive high-quality training data, and the shortage of such data has become the main performance bottleneck. In practice, data augmentation is often used as a countermeasure to mitigate this issue, by synthesizing additional training data based on existing ones. However, most existing practice of data augmentation in source code learning limits simple code refactoring methods and is not sufficiently effective. In this work, in light of the graph nature of source code, we propose to apply the data augmentation methods used for graph-structured data in graph learning to the tasks of source code learning, and we conduct a comprehensive empirical study to evaluate whether such new ways of data augmentation are more effective than the existing simple code refactoring methods. Specifically, we evaluate 4 critical software engineering tasks and 7 neural network architectures to assess the effectiveness of 5 data augmentation methods. Experimental results identify that, compared to other methods, Manifold-Mixup can greatly improve the accuracy of the trained models for source code learning.

Read the paper · More papers on PaperTik