CG-Bench: Can Language Models Assist Call Graph Construction in the Real World?
Ting Yuan, Wenrui Zhang, Dong Chen, Jie Wang · 2025
Language models for coding are shifting their focus from function-level to repository-level, with complex function invocations. We introduce CG-Bench, the first manually constructed benchmark that measures the ability to understand call graphs for language models. This benchmark contains 104 call sites and related code snippets associated with call chains from 7 representative open-source C/C++ projects. Language models are tasked with inference the calling targets from them. We evaluated four popular language models on CG-Bench. Surprisingly, all four models with different prompt settings achieve accuracy greater than 50% and Deepseek-6.7b with few-shot prompts reaches 69.70%. We further show four findings from a micro study, which demonstrates that using language models for call graph construction is promising and the performance can be improved by prompt hacking, removing irrelevant information, etc.