Why and How to Control Cloning in Software Artifacts
Elmar Juergens · mediaTUM – the media and publications repository of the Technical University Munich (Technical University Munich) · 2011
The majority of the total life cycle costs of long-lived software arises after its first release, during software maintenance.Cloning, the duplication of parts of software artifacts, hinders maintenance: it increases size, and thus effort for activities such as inspections and impact analysis.Changes need to be performed to all clones, instead of to a single location only, thus increasing effort.If individual clones are forgotten during a modification, the resulting inconsistencies can threaten program correctness.Cloning is thus a quality defect.The software engineering community has recognized the negative consequences of cloning over a decade ago.Nevertheless, it abounds in practice-across artifacts, organizations and domains.Cloning thrives, since its control is not part of software engineering practice.We are convinced that this has two principal reasons: first, the significance of cloning is not well understood.We do not know the extent of cloning across different artifact types and the quantitative impact it has on program correctness and maintenance efforts.Consequently, we do not know the importance of clone control.Second, no comprehensive method exists that guides practitioners through tailoring and organizational change management required to establish successful clone control.Lacking both a quantitative understanding of its harmfulness and comprehensive methods for its control, cloning is likely to be neglected in practice.»A man's gotta do what a man's gotta do« Fred MacMurray in The Rains of Ranchipur »A man's gotta do what a man's gotta do« Gary Cooper in High Noon »A man's gotta do what a man's gotta do« George Jetson in The Jetsons »A man's gotta do what a man's gotta do« John Cleese in Monty Python's Guide to LifeTo operationalize clone control, comprehensive tool support is required that supports all of its steps.Existing tools, however, typically focus on individual aspects, such as clone detection or change propagation, or are limited to source code and thus cannot be applied to specifications or models.Furthermore, most detection approaches are not both incremental and scalable.They thus cannot provide real-time results for large evolving software artifacts.Dedicated tool support is thus required for clone control.Problem We need a better understanding of the quantitative impact of cloning on software engineering and a comprehensive method and tool support for clone control.