IDENTIFICATION OF SOFTWARE CLONE FILES USING MACHINE LEARNING
Mediboyina Venkata Swapna · INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT · 2023
One technique to reduce the likelihood of introducing a bug is to identify code clone files. Refactoring, often known as removal, is the process of making software more understandable and maintainable. When additional clones are discovered in software programs, we need a method to help developers with refactored code and to enhance software quality. This method must tell developers which clone files require reworking. To provide developers with recommendations on which files require code refactoring, our research proposes a novel learning technique that automatically extracts features from the identified code clones and trains models to present a novel technique to enhance classification performance by transforming outliers into Unknown clone sets. The Eclipse dataset, which comprises 213 Java software files, was utilized for this project. Support Vector Machine (SVM), Random Forest, Bagging, and K-Nearest Neighbors (KNN) are the algorithms that are employed. We compare these four categorization models and recommend the model with the highest accuracy. Our tool can be developed and applied to reduce system bugs. By recognizing software clone files and eliminating duplication code, our study can enhance clone maintenance. Also, the possibility of bad design for a system, difficulty in a system improvement or modification, and introduction of a new bug can be decreased by identifying and refactoring clones. KEYWORDS: Cloning, Refactoring, Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Decision trees.