Explainable Multi-Agent LLM Framework for Phishing Email Detection via Role-Specialized Evidence Decomposition

Tanya Yadav, Mohammad Masum · Electronics · 2026

Phishing email remains a persistent and operationally critical cybersecurity threat, yet existing detection approaches, including traditional machine learning and single-pass large language model systems, either lack native interpretability or provide explanations that are difficult to standardize and audit. This paper introduces an explainable multi-agent LLM framework that decomposes phishing evidence across three role-specialized agents focused on linguistic patterns, psychological manipulation, and sender identity consistency. The framework then aggregates specialist outputs through schema-governed synthesis, enabling each intermediate and final decision to be structured, comparable, and auditable. The central contribution is the treatment of role-specialized evidence decomposition and explanation structure as first-class design constraints rather than post hoc additions. The framework is evaluated on a fixed 1000-email subset drawn from a unified TREC/Nazario corpus of 56,212 emails under controlled zero-shot conditions. The full multi-agent Meta-Judge system achieves Macro-F1 = 98.28% and phishing recall = 99.45%, improving Macro-F1 by 6.3 percentage points over a zero-shot single-model GPT-4o-mini baseline. Paired statistical testing confirms that this improvement is significant and is driven primarily by reduced false positives on legitimate emails while preserving high phishing recall. Additional evaluation on an independent LLM-attributed email benchmark shows a consistent Macro-F1 improvement of 0.0773 over the zero-shot baseline under distribution shift. Ablation results show that role-specialized decomposition is the primary performance driver, while deterministic voting provides a competitive raw-classification aggregator and Meta-Judge synthesis provides structured, analyst-facing explanations. These results indicate that role-specialized evidence decomposition combined with schema-governed explanation can improve both detection reliability and auditability in phishing classification workflows.

Read the paper · More papers on PaperTik