QAEval: Mixture of Evaluators for Question-Answering Task Evaluation
Tan Yue, Rui Mao, Xuzhao Shi, Shuo Zhan, Zuhao Yang, Dongyan Zhao · 2025
Question answering (QA) tasks serve as a key benchmark for evaluating generation systems.Traditional rule-based metrics, such as accuracy and relaxed-accuracy, struggle with openended and unstructured responses.LLM-based evaluation methods offer greater flexibility but suffer from sensitivity to instructions, robustness issues, and high computational costs.To overcome these challenges, we introduce QAEval, a hybrid framework combining rule-based reliability with LLM-based adaptability.QAEval utilizes two high-quality datasets: QAExtract for short-answer extraction and QAScore for scoring model training.By integrating a Mixture of Evaluators model with Dynamic Load Balancing Optimization, QAEval enables accurate, cost-effective QA evaluation.Experimental results show it outperforms models like GPT-4o and Claude-3, achieving 92.3% accuracy with only 0.6B parameters.