MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors

Baraa Hikal, Mohamed Basem, Islam Oshallah, Ali Hamdi · 2025

We present MSA-MATHEVAL, our submission to the BEA 2025 Shared Task on evaluating AI tutor responses across four instructional dimensions: Mistake Identification, Mistake Location, Providing Guidance, and Actionability.Our approach uses a unified training pipeline to fine-tune a single instructiontuned language model across all tracks, without any task-specific architecture modifications.To improve prediction reliability, we introduce a disagreement-aware ensemble inference strategy that enhances coverage of minority labels.Our system achieves strong performance across all tracks, ranking 1 st in Providing Guidance, 3 rd in Actionability, and 4 th in both Mistake Identification and Mistake Location.These results demonstrate the effectiveness of scalable instruction tuning and disagreementdriven modeling for robust, multi-dimensional evaluation of LLMs as educational tutors.

Read the paper · More papers on PaperTik