Generated code in studies on clone rates
Rainer Koschke, Moritz Weinig · 2018
Various earlier studies have measured clone rates for diverse projects. One of the reasons for exceptionally high clone rates for individual source files was found to be auto-generated code. Automatically generated code is generally not maintained and, hence, should be excluded from clone-rate measurements. This kind of code might even introduce a bias to clone rates of projects when there is a large amount of generated code and clone rates for generated files generally deviate from the average clone rate for handwritten code. While some generated files stuck out with clone rates above the average in earlier studies, we do not know whether this is generally the case and how much code is actually generated automatically. This paper investigates the amount of generated files in projects, whether clone rates for generated files really differ from handwritten code, and - overall - whether generated code in fact introduces a bias to clone rates. We heuristically detect generated files in a very large open-source project corpus of programs written in C, C++, C#, or Java and report the number of projects with generated code. For these projects, we compare clone rates of generated and handwritten files. Our results show higher clone rates for generated files. Moreover, when we aggregate clone rates from files to projects, the clone rates of projects with at least one generated file are also slightly higher than in projects for which no generated files were detected. Our results suggest that researchers should indeed take special care to exclude generated code in studies on clone rates.