Automating Benchmark Generation for LLMs in Software Engineering: Challenges and Opportunities
Nimrod Busany, Hananel Hadad, Gil Rosenblum, Ramya Ramachandran, Zofia Maszlanka, Rohit Shashank Shelke, Okhaide Akhigbe, Daniel J. Amyot · 2025
As Large Language Models (LLMs) become increasingly integral to software engineering tasks, the need for extensive evaluation benchmarks grows. Yet, creating these benchmarks manually is costly and time-consuming, posing a significant barrier to the effective testing and deployment of LLM-based systems. In this position paper, we highlight the challenges associated with manual benchmark creation using LLMs and discuss the potential of using LLMs themselves to automate this process. We explore critical issues such as ensuring dataset quality and comparability to human-crafted benchmarks, assisting users in validating generated examples, reducing the need for manual feedback, lowering associated costs, and facilitating customization across diverse tasks. By examining these challenges and sharing insights from our preliminary tool development, we aim to spark discussion and collaboration within the community to develop scalable solutions for benchmark generation.