PPTC-R benchmark: Towards Evaluating the Robustness of Large Language Models for PowerPoint Task Completion

Ze Zhang, Yiduo Guo, Yaobo Liang, Dongyan Zhao, Nan Duan · 2024

The growing dependence on Large Language Models (LLMs) for completing user instructions necessitates a comprehensive understanding of their robustness in real-world situations.To address this need, we introduce the Power-Point Task Completion-Robustness (PPTC-R) benchmark, designed to evaluate LLMs' robustness to user task instructions and different software versions (PowerPoint versions).Specifically, we create adversarial user instructions by manipulating instructions at the sentence, semantic, and language levels.To assess robustness across software versions, we vary the number of available APIs to simulate both the latest and earlier version environments.We benchmark 3 closed-source and 4 open-source LLMs against these robustness settings to evaluate how variations affect their API calls for task completion.Our findings reveal that GPT-4 demonstrates the highest performance and strong robustness, especially in version updates and multilingual settings.However, all LLMs show a significant decline in robustness when faced with multiple simultaneous challenges (e.g., multi-turn interactions), resulting in notable performance drops.We further analyze the robustness behavior and error patterns of LLMs in our benchmark, providing valuable insights for researchers to understand LLM robustness in task completion and to develop more resilient LLMs and agents.The code and data are available at https: //github.com/ZekaiGalaxy/PPTCR. * We put the new APIs in the supplementary

Read the paper · More papers on PaperTik