Evaluating Large Language Models for Code Generation: Assessing Accuracy, Quality, and Performance

Mohammed A. Shehab, Mohammad Wardat, Safwan Omari, Yaser Jararweh · 2024

Large Language Models (LLMs) are increasingly utilized for software engineering tasks, including code generation. While prior research has primarily focused on code completion, there remains a gap in understanding LLMs’ ability to generate code from Natural Language Processing (NLP) descriptions. In this study, we investigated three LLMs’ performance in generating code from scratch based on NLP task descriptions. We employed three evaluation levels, Accuracy, Quality, and Performance, to assess the LLMs’ results. Accuracy metrics focused on error types and counts, quality assessed code readability and maintainability, and performance measured the models’ ability to generate optimized solutions. Codex, Copilot, and PaLM2 exhibited error rates of 9.68%, 32.25%, and 48.38%, respectively. Copilot showed the best performance with a higher maintainability index. However, all LLMs struggled with tasks of high complexity, with time complexities of O(m•n•log(n)), as evidenced by their performance metrics.

Read the paper · More papers on PaperTik