NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors
Numaan Naeem, Sarfraz Ahmad, Momina Ahsan, Iqbal Hasan · 2025
MISTAKE IDENTIFICATION in the BEA 2025 SHARED TASK ON PEDAGOGICAL ABILITY ASSESSMENT OF AI-POWERED TU-TORS.The task involves evaluating whether a tutor's response correctly identifies a mistake in a student's mathematical reasoning.We explore four approaches: (1) an ensemble of machine learning models over pooled token embeddings from multiple pretrained langauge models (LMs); (2) a frozen sentencetransformer using [CLS] embeddings with an MLP classifier; (3) a history-aware model with multi-head attention between token-level history and response embeddings; and (4) a retrieval-augmented few-shot prompting system with a large language model (LLM) i.e.GPT 4O.Our final system retrieves semantically similar examples, constructs structured prompts, and uses schema-guided output parsing to produce interpretable predictions.It outperforms all baselines, demonstrating the effectiveness of combining example-driven prompting with LLM reasoning for pedagogical feedback assessment.Our code is available at https://github.com/NaumanNaeem/BEA_ 2025.