Poster: Benchmarking of Code Generative LLMs
Mirza Masfiqur Rahman, Ashish Kundu, Elisa Bertino · 2024
Generative LLMs have proven to be valuable code generators, thus enabling code copilots and meeting several requirements in software engineering. However, several questions arise: How good an LLM is as a software engineer? How secure is the code generated or fixed by an LLM? These are complex questions to address; however, addressing them is critical for enabling trustworthy software development ecosystems. Addressing those questions requires a designing rigorous benchmark to evaluate: (i) the code that is generated/completed, and (ii) the generative LLMs themselves. In this paper, we propose an automated benchmarking system covering the different aspects of the generated code and the LLMs that generate the code. We also propose the concept of a benchmark dependency graph, coupled with an automated benchmark scoring process that provides a vector of scores on how “good” or how “bad” an LLM is, or the code generated/modified by it is, additionally the artifacts generated or modified around the code.