Rethinking Distractor Quality in Multimodal Multiple-Choice Questions: Automated Evaluation and Hard Benchmark Construction
Wenjian Ding, Yao Zhang, Jun Wang, Zhenglu Yang · Computers · 2026
In Multimodal Multiple-Choice Questions, distractors play a pivotal role in rigorously evaluating the cross-modal reasoning capabilities of Multimodal Large Language Models by serving as plausible yet incorrect options. A comprehensive and reliable evaluation of distractor quality is therefore imperative for fostering genuine progress in this domain. However, prevailing evaluation approaches face a fundamental dilemma: they either rely on model-based metrics that fail to fully capture semantic nuances, or depend on human evaluation, which is resource-intensive and prone to subjective bias. To address these limitations, we introduce a comprehensive suite of 9 automated metrics, spanning both intrinsic and extrinsic dimensions, to reliably quantify distractor quality. Leveraging this framework, we propose a metric-driven ensemble strategy for constructing hard benchmarks. Specifically, we aggregate candidate pools from diverse advanced baselines and rigorously select the optimal subset of distractors that yield the highest quality scores under our proposed metrics. Extensive evaluations involving 33 Multimodal Large Language Models across 16 diverse benchmarks demonstrate that our method generates distractors with significantly higher confusability, posing a more rigorous challenge to current state-of-the-art models.