Advancing Agile Software Cost Estimation Through Data Synthesis: A Comparative Analysis of Five Generation Techniques
Xiaoyan Zhao, Zulkefli Mansor, Rozilawati Razali, Mohd Zakree Ahmad Nazri, Xin Xiong · IEEE Access · 2025
Agile has been used in software development for over 20 years and is the preferred development method for more than 85% of software companies. However, cost estimation in agile development remains a significant challenge. This is reflected in the fact that the accuracy of estimation still needs to be improved, and most cost estimation techniques still depend on the experience and knowledge of the team. While machine learning algorithms have performed better in this area, the lack of sufficient agile cost data hinders large-scale training and in-depth research. To address this issue, this study selected five data generation techniques—Variational Autoencoder (VAE), Wasserstein Generative Adversarial Network (WGAN), Synthetic Minority Over-sampling Technique for Nominal and Continuous Features (SMOTE-NC), Data Augmentation for Tabular Data (Augmentation), and Tabular Data Diffusion Probabilistic Models (TabDDPM)—based on the characteristics of agile cost data. Using cost data from 75 agile projects, these techniques were employed to generate three sets of data with sizes of 200, 500, and 1000. A performance evaluation model was created based on consistency, authenticity, diversity, and effectiveness to verify the performance of these generated data. The experimental results show that WGAN consistently scored 16 out of 20 points across all three data sets, excelling in data consistency and authenticity. SMOTE-NC and Augmentation followed. SMOTE-NC scored 15 out of 20 points for all data sizes and performed best in effectiveness, with an MMRE of 88.16% and a PRED (0.2) of 84.5%. Augmentation had the highest performance when generating 1000 data points. These findings highlight the potential of data generation technologies, especially WGAN, in improving agile cost estimation and providing guidance on selecting appropriate data amounts. This lays a foundation for further development of machine learning algorithms in this field and offers valuable insights for other researchers.