ConfidenceCal: Enhancing LLMs Reliability through Confidence Calibration in Multi-Agent Debate

Yilin Bai · 2024

Multi-agent debate enhances large language models (LLMs) by facilitating collaborative interactions that improves problem solving and decision making. Despite the potential of LLMs, they often generate seemingly confident answers. This can lead to communication barriers and poor decision-making due to the lack of effective mechanisms for expressing and adjusting confidence levels. To address this issue, we developed the ConfidenceCal framework, which aims to improve the reliability of LLMs by incorporating calibrated confidence levels into multi-agent debate. In addition, we explore the use of textual prompts to convey confidence, and also adapt the attention mechanism to adjust token weights based on confidence levels. Extensive evaluation of multiple benchmarks has shown that Confidence Cal significantly reduces misleading confidence and increases the trustworthiness of multi-agent communication.

Read the paper · More papers on PaperTik