DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues

Kyochul Jang, Donghyeon Lee, Kyusik Kim, Dongseok Heo, Taewhoo Lee, Woojeong Kim, Bongwon Suh · 2025

Existing function-calling benchmarks focus on single-turn interactions.However, they overlook the complexity of real-world scenarios.To quantify how existing benchmarks address practical applications, we introduce DICE-SCORE, a metric that evaluates the dispersion of toolrelated information such as function name and parameter values throughout the dialogue.Analyzing existing benchmarks through DICE-SCORE reveals notably low scores, highlighting the need for more realistic scenarios.To address this gap, we present DICE-BENCH, a framework that constructs practical functioncalling datasets by synthesizing conversations through a tool graph that maintains dependencies across rounds and a multi-agent system with distinct personas to enhance dialogue naturalness.The final dataset comprises 1,607 high-DICE-SCORE instances.Our experiments on 19 LLMs with DICE-BENCH show that significant advances are still required before such models can be deployed effectively in realworld settings.Our code 1 , and data 2 are all publicly available.

Read the paper · More papers on PaperTik