A Similarity-based Stacked Deep Learning Architectures for Detection of Software Clones
Abdullah M. Sheneamer · 2024
Reducing duplication and preserving code quality require the detection of software clones, or similar or identical code parts. The accuracy and scalability of conventional clone detection techniques are threatened by the growing complexity and scale of software projects. This study presents a novel formal similarity model that includes various similarity measures to evaluate syntactic and semantic distances across method blocks. A similarity-based stacking deep learning architecture intended to improve software clone detection. Our work is unique in that it uses a variety of similarity scores and measurements as features in layered deep learning to detect code clones based on semantic code features. Our results show that our approach significantly outperforms existing code clone detectors, achieving a 99% success rate in detecting cloned code based on F-score, recall, and precision, and 98-100% accuracy in most cases.