Measuring Efficient Code Generation with GEC
Yue Pan, Chen Lyu · 2023
Although efficiency is one of the core metrics in programming, recent large-scale language models often face the issue of “inefficient code” generation, which struggles to meet the real-time requirements of algorithms. However, there is relatively little research on evaluating the selection of efficient algorithms, and it is not easy to rigorously assess a model’s ability to correctly choose efficient algorithm solutions. Furthermore, the selection of efficient algorithm solutions often relies on the appropriate application of problem-solving skills, necessitating more in-depth research on algorithm reasoning. To address this challenge, we introduce the Generation of Efficient Code (GEC) benchmark, which aims to evaluate the ability to select efficient algorithm solutions. Unlike code generation, our benchmark focuses on a model’s ability to generate satisfactory efficient code when given a natural language description and inefficient code. We propose two novel metrics to examine the efficiency of the generated code and assess the model’s ability to generate efficient code. Our benchmark includes 3,712 problems, 31,577 combinations of efficient and inefficient code pairs, and 13,092 alternative efficient codes. We evaluate the performance of mainstream code generation models on the GEC benchmark. As the societal importance of code efficiency increases in the coming years, our benchmark will provide an essential measurement standard for tracking research progress. Our dataset and models are open-source and can be accessed at https://github.com/CodeGeneration2/Efficient-Code-Generation-with-GEC.