Data augmentation of Python code refactoring datasets based on LLMs

V. MOLDOVAN, Rares Patcas, Simona Motogna · Journal of Systems and Software · 2026

Refactoring is a crucial software engineering practice aimed at improving code quality, yet automatically detecting and predicting refactoring activities remains difficult due to the limited availability of labeled data. This study examines how data augmentation can strengthen refactoring detection models. We experimented with 4 augmentation methods, namely Back Translation from textual description, Back Translation from Java, Generating Python code from existing comments, and Generating Python code from generated comments, and applied each of them using 4 large language models, specifically Gemini, GPT-4o, DeepSeek, and GPT-5-Nano. Through these combinations, we generated new instances of paired pre-refactor and post-refactor functions derived directly from the original ground truth pairs. We compute embedding-based cosine similarities between original functions, refactored functions, and their generated counterparts to evaluate semantic fidelity and proximity. We also validate and balance the resulting dataset to ensure it remains suitable for downstream machine learning tasks. Our findings suggest that many augmented samples remain semantically aligned with the original functions while adding diversity that improves robustness, reduces overfitting, and enhances generalization in automated refactoring detection.

Read the paper · More papers on PaperTik