Code clones detection using machine learning technique: Support vector machine
Shruti Jadon · 2016
Code clones defined as sequence of source code that occur more than once in the same program or across different programs are undesirable as they increase the size of program and creates the problems of redundancy. Fixing of bugs detected in one clone require detection of all clones. Hence, it is imperative to identify and remove all code clones in a program. The focus of previous research work on the code clone detection was to find identical clones, or clones that are identical up to identifiers and literal values. But, detection of similar clones is often important. In the present paper it is proposed to generate the feature sets after parsing the given C program for code fragments and then match their similarity. On the basis of feature sets the classification of algorithm is being performed by using the Support Vector Machine (SVM) as a machine learning tool. The output of the machine tool would be the similarity ratio with which the two C programs are related to each other and also the class in which they would occur. It was observed that the test results of the tool implementation show detection of code clones in the program and its accuracy increases with the increase in number of instances.