Classifying and mining design diagrams in source code repositories using transfer learning

Sergio Rodríguez, Luis Daniel Benavides Navarro, Héctor Cadavid, Wilmer Garzón · Array · 2025

This paper reports on the implementation of a design diagram classifier using transfer learning and fine-tuning. For the study, we built a labeled data set with 5981 images identifying the following design diagram categories: none (no design diagram represented), activity diagram, sequence diagram, class diagram, component diagram, use case diagram, and cloud diagrams. We then used the dataset, transfer learning, and fine-tuning techniques to train a specialized DenseNet169 convolutional network, starting from a model pre-trained on ImageNet. The newly trained network achieved a prediction accuracy of 98.6% ( 0 . 986 ± 0 . 0017 ) and an F1-score of 98.3% ( 0 . 983 ± 0 . 0025 ). Repository mining techniques were then used to analyze 2,469,206 images, equivalent to 231 GB of data, obtained from 287,201 repositories. The analysis revealed that design diagrams are often outdated relative to the project’s evolution; on average, diagrams were last updated 554 days prior to the repository’s latest commit. This significant lag suggests that while diagrams are present, practitioners primarily rely on the source code as the living design artifact.

Read the paper · More papers on PaperTik