Research of Evaluating the Effectiveness of Large Language Models in Identifying and Correcting Code Anomalies
Tianyi Chen · 2024
Large language models (LLMs) are powerful tools that can generate code in various programming languages, and assist programmers in writing, debugging, or improving code. However, LLMs can also introduce or overlook code anomalies, such as errors, bugs, vulnerabilities, or deviations from the expected behavior or style of a program, due to their lack of semantic understanding, generalizability, or reliability. In this paper, we investigate the use of anomaly detection techniques to evaluate the effectiveness of LLMs, such as OpenAI Codex, in identifying and correcting code anomalies. We propose a novel methodology for evaluating the effectiveness of LLMs in identifying and correcting code anomalies, using anomaly detection techniques, and we apply it to several LLMs, such as OpenAI Codex. We present and analyze the results of the evaluation, and compare the performance of different LLMs, anomaly detection methods, and evaluation metrics. We also discuss the implications, challenges, and limitations of using LLMs for code quality and security, and suggest some directions for future research. We hope that this paper will inspire and inform researchers, developers, and users of LLMs and code quality and security, and that it will stimulate further research and innovation in this field.