Towards Affective Evaluation of STEM Education: Leveraging MLLMs in Project-Based Learning
X. H. Wu, Yanhao Jia, Qinglin Zhang, Yiran Qin, Luwei Xiao, Shuai Zhao · IEEE Transactions on Affective Computing · 2026
Project-Based Learning (PBL) involves a variety of highly correlated multimodal data, making it a vital educational approach within STEM disciplines. With the rapid development of multimodal large language models (MLLMs), researchers have begun exploring their potential to enhance tasks such as information retrieval, knowledge comprehension, and data generation in educational settings. However, existing benchmarks lack both a flexible free-form output structure and rigorous human expert validation, resulting in limited affective engagement and insufficient capacity to evaluate real-world educational tasks. Additionally, few methods have developed automated pipelines to assist with the complex responsibilities of teachers leveraging MLLMs, largely due to model hallucination and instability, which further limit their affective responsiveness and result in unreliable implementation. To address this gap, we introduce Affective-PBLBench, an affect-informed benchmark designed to evaluate complex reasoning grounded in domain-specific knowledge and long-context understanding, thereby challenging models with tasks that closely resemble those handled by human experts. Recognizing that affective framing is pivotal for holistic student assessment, Affective-PBLBench specifically includes a branch for evaluating students' affective states, facilitating the monitoring and preliminary screening of students' mental health conditions. We also build a new dataset, PBL-STEM, for this complex scenario, which contains over 500 projects with different modalities and multi-disciplinary contexts. To establish reliable ground truth, we adopt the Analytic Hierarchy Process (AHP), utilizing expert-driven pairwise comparisons to derive structured and weighted evaluation criteria. We assess the performance of 15 leading MLLMs/LLMs using Affective-PBLBench and demonstrate that even the most advanced models achieve only 59% rank accuracy, underscoring the significant challenges presented by this benchmark. We believe Affective-PBLBench will serve as a catalyst for the development of more capable AI agents, ultimately aiming to alleviate teacher workload and facilitating a holistic understanding of students' mental health.