An improved method for tree-based clone detection in Web Applications
Chaoqun Li, Jianhua Sun, Hao Chen · 2014
Clone detection has been an active area for decades and many tools have been proposed. Existing researches show that in traditional software clones achieve to 13%-20%, and the clone rate in Web Application area may be higher. In this paper, we propose an improved method for code clone detection. Our approach is based on the randomized kd-trees with dimensionality reduction to cluster the characteristic vectors. We have implemented our algorithm and tested it on large software projects written in Java and PHP including JDK and 7 Web applications. Our experimental results show that our tool is efficiency both in speed and accuracy. In addition, we conducted empirical studies on Web applications and find that the cloned rates achieve to 5-82% in PHP Web Applications.