Prevalence and Prediction of Unseen Co-Changes: A Graph-Based Approach
Amit Kumar, Hrishikesh Ethari, Yugandhar Deasi, Sonali Agarwal · 2024
Co-changes refer to the phenomenon wherein two or more software entities are modified together within the same commit or when they are changed to accomplish a specific task or functionality. The key to accurate co-change prediction lies in effectively predicting unseen co-changes-those occurring between entities that have not been co-changed before. However, despite considerable research on co-change patterns and prediction, there remains a significant gap in understanding unseen co-changes, including their prevalence, complexity, and predictability. We model co-changes as a graph, treating unseen co-change prediction as a link prediction task. Our method leverages file proximity measures derived from both homogeneous and heterogeneous networks, alongside other similarity measures, to predict these unseen co-changes. Analysis of 14 Apache Software Foundation projects revealed a significantly higher prevalence of unseen co-changes (up to 23x more frequent in specific projects and 7x on average) compared to recurrent co-changes. Interestingly, comparisons of co-change complexity based on file distance in the directory structure revealed no decisive differences between the two types. Our graph-based approach achieved good accuracy in predicting unseen co-changes (average AUROC of 0.84, with some projects reaching up to 0.98). Our method achieved significantly better performance than the baseline approach, demonstrating an average recall of 90% and a precision of 38%. While the precision value might seem modest, our approach achieves very high precision@k values (near 100% for 13 out of 14 projects up to$k=100$), underlining its effectiveness in real-world applications.