Characteristics and Automated Detection and Refactoring of Data Clumps
Nils Baumgartner · osnaDocs (Osnabrück University) · 2026
In object-oriented (OO) systems, such as source code projects and Unified Modeling Language (UML) class diagrams, maintainability can be negatively affected by so-called bad smells, which increase development costs. These bad smells are not faults directly but indicate poor development practices. Refactoring improves internal structures and removes such smells. Data clumps, which are recurring groups of variables, represent a specific type of these bad smells and are among the most commonly found in software projects. However, despite their prevalence, their characteristics remain largely unexplored, and tools and datasets for effective detection and refactoring are lacking. This thesis contributes to close this gap by improving the understanding of data clumps and providing a live detection and refactoring framework that integrates seamlessly into the developer workflow. The developed approach introduces a method for live detection based on structural features of data clumps. At its core, it applies information retrieval techniques that leverage an inverted index dictionary to efficiently identify recurring variable groups across an analyzed project. In addition, a semi-automatic refactoring support system to eliminate data clumps is developed. By means of these tools a publicly available dataset has been created as a base for further research work on data clumps. The dataset includes more than 8 million data clumps with detailed, line-level granularity across 23 open-source projects and 3,290 timestamps, as well as more than 100,000 UML class diagrams. In some projects, this longitudinal dataset spans up to 27 years of development history. The main contributions of this thesis include novel concepts for the live detection, prioritization, and refactoring of data clumps. In the analyzed source code projects, data clumps occur predominantly in method parameters, while in UML class diagrams, they appear almost exclusively in class attributes. Data clumps occur mostly locally within individual classes, and their number increases over the lifetime of a software project. A strong positive correlation between the presence of data clumps and fault rates is identified in many projects. These findings support the integration of real-time assistance for refactoring to reduce maintenance costs and improve software quality. The dataset provides a foundation for future research and benchmarking of data clump detection tools. The results also highlight the potential for prioritization and emphasize the need for further research in combination with other bad smells, visualization techniques, and additional programming languages.