Detection and evolution analysis of code clones for efficient management of large-scale software systems
Eunjong Choi · Institutional Repositories DataBase (IRDB) · 2015
In recent decades, large-scale software systems have become mainstream.Such software systems have complicated the maintenance process by increasing efforts such as inspection and understanding of the existing source code.Therefore, to maintain these systems, a great deal of work and time are necessary.To alleviate this problem, this research focus on a well-known factor hindering the software maintenance task, a code clone (i.e., a code fragment that has other code fragments identical or similar to it in the source code).It is widely believed that code clones complicate software maintenance.For example, when changes to code clones in a clone set (i.e., a set of code clones that are identical or similar to each other) are inconsistent, the developer needs to identify inconsistently changed code clones and apply consistent changes to them.Thus far, many tools and techniques have been proposed for supporting the detection and management of code clones.However, most are insufficient for supporting code clone related tasks during the software maintenance process for large-scale software systems.To resolve this problem, this study attempts to solve two important problems that code clones face.That is, "Which type of normalization dose make code clones to detected with high speed from large-scale software systems?" and "Which supports are necessary for more widely used tools that support clone refactoring?".To solve the first problem, this research proposes six approaches for detecting code clones with preprocessing input source files using different degrees of normalizations (e.g., the removal of white spaces, tokenization, and the regularization of identifiers).More precisely, each type of normalization is applied to the input source files, and equivalence class partitioning of the files is then conducted during the preprocessing.Code clones are then detected from a set of files that are representatives of each equivalence class using a token-based code clone detection tool called CCFinder.The proposed approaches can be categorized into two types, an approach with non-normalization and approaches with normalization.The former type is the detection of only identical files without normalization, whereas the latter category is the detection of identical files with different degrees of normalization iii First of all, I would like to express my most sincere gratitude to my respected supervisor Katsuro Inoue.Without his warm supports and valuable comments regarding my research, this study would not have been possible.I feel extremely happy and blessed for the opportunity to have been supervised by him and have been inspired by his enthusiasm and integral view toward research.I would especially like to express my gratitude to Professors Toshimitsu Masuzawa, and Shinji Kusumoto for their valuable comments and helpful suggestions regarding this thesis.