Auto-Generated Code Detector - A Preprocessor to Enhance Security, Quality, Maintenance and Accurate Productivity Calculation of Code

Ratnesh Parihar, Priya Kar, Shravan Kumar · 2021 7th International Conference on Computer and Communications (ICCC) · 2021

Auto-generated code detection is emerging fast as a must-have pre-processing activity for software companies. If unchecked, auto-generated code can trigger security issues and incorrect quality and productivity calculations. Besides, devising a generic algorithm is difficult because of the colossal size of repositories and diversified technology stack. That is why researchers are trying to detect, scrutinize, verify, and validate code in terms of security, copyrights, and code licenses to exclude auto-generated code during quality and productivity analysis. It will improve the software reliability and integrity, code quality, maintainability of deliverable, correct productivity calculation and ensure code’s adherence to copyrights and licenses.Auto-generated code detection detects machine-generated code, external libraries, minified files/packages/bundles files, and anomaly files [1]. Among these, the anomaly files have one-time large commits and code from other repositories. Such files force the quality analysis tools to slow down or hang or produce incorrect quality reports causing major bottlenecks. Also, the manual identification of auto-generated code is time-consuming and not sustainable. These acted as a catalyst for our extensive research to detect auto-generated code with high accuracy. We used our domain expertise and machine learning skills to write algorithms for better detection. As its outcome, our proposed technique can identify whether the source code is auto-generated or not by mining the code and collecting data from version control. We have run our technology-agnostic algorithm across various repositories within Talentica as well as for external repositories belonging to Google, Facebook, Twitter, Thoughtworks, Netflix and have achieved an average of 96.85% accuracy.

Read the paper · More papers on PaperTik