Cloned Code Clustering for the Software Product Line Engineering Approach to Developing a Family of Products
Tae-Young Kim, Jihyun Lee, Sungwon Kang · 2024
Cloned code clustering identifies identical or similar code fragments, i.e., cloned code, from products of a product family, and then constructs cloned code clusters, a set of files, each of which contains the same cloned code. It is an essential step for migrating from the Clone-And-Own (CAO) approach to the Software Product Line Engineering approach for developing a family of products. This paper proposes a method for cloned code clustering based on files that share cloned code. The method proposed in this paper identifies clusters to use for constructing a product line platform at the source code level in such a way that it works regardless of cloning-in-the-small or cloning-in-the-large; it does not need to know what code is original or what code is cloned; and its results remain consistent regardless of the comparison order or the relative similarity of products of a product family. For evaluation we applied our method to Apo-Games and ArgoUML developed with the CAO approach and confirmed that our method correctly constructs code clusters.