Contemporary Software Cloning Detection Methods: Meta-Analyses
Chavi Ralhan, Navneet Malik · 2024
This study examines numerous open-source clone detection studies. Code tokens feed most ML and DL algorithms. Duplicate code algorithms have nine categories (XX). Lexical tokens first. Distance-based clone detection follows. Third, tree-based and semantic clone detection Researchers combine detection approaches to generate the sixth category. This study shows that academics employ machine learning and deep learning to find textual clones. Unsupervised and supervised machine learning can find clones. K-means and DBScan are popular machine learning approaches. Naive Bayes, Decision Tree, SVM, and others are prominent supervised algorithms. Datamining. Subgraph and FP-growth find code. Applications find duplicate codes. Visual Studio or GitHub apps Clone detection checks for vulnerability, licensing, and security in several apps. Code rearrangement and analysis tools find duplicate code chunks. Academically and commercially, it’s underdeveloped. Analysis The largest code repository has duplicate Java, C++, C, Python, and JavaScript code, and most non-focked projects have borrowed files from other projects (inter-project level analysis). Some researchers want faster, smaller clone-detecting technologies.