Adversarial-Resistant AI Tutoring: A Dual-LLM Architecture for Mitigating H-CoT Attacks and Ensuring Pedagogically-Aligned Code Assistance
Khan Tahsin Abrar · 2025
Current AI tutoring systems exhibit systematic pedagogical misalignment, with students successfully bypassing educational guardrails through adversarial prompting techniques. Through empirical analysis of Harvard's CS50 Duck AI, we document critical vulnerabilities including emotional manipulation exploits that trigger inappropriate code generation, undermining learning objectives. Despite sophisticated fine-tuning approaches, CS50 Duck AI exhibits persistent instruction dilution with 22% of responses containing inappropriate code blocks. We propose a dual-LLM architecture employing architectural separation between reasoning and compliance validation, combined with behavioral design elements including strategic response delays and adversarial training protocols. Our system utilizes a primary DeepSeek Coder model for educational content generation and a specialized 1B parameter trimmer model for pedagogical compliance verification. Experimental evaluation across 50 educational scenarios demonstrates 86% reduction in inappropriate code provision compared to baseline systems, while maintaining pedagogical effectiveness as measured by student engagement metrics and learning outcome alignment. The proposed approach addresses fundamental limitations in current educational AI systems through systematic behavioral design and architectural robustness mechanisms.