Agentic AI system for generating evidence-based second medical opinions
Diana Hawashin, Khaled Salah, Raja Jayaraman, Samer H. Ellahham · Clinical eHealth · 2026
Second medical opinions play a critical role in reducing diagnostic uncertainty and supporting high-stakes clinical decision-making. However, healthcare systems often struggle to provide them in a timely manner because of complex workflows and the need to interpret evolving evidence-based guidelines. Many AI-based clinical decision support systems (CDSS) remain restricted to narrow medical domains or rely on static knowledge repositories, limiting their ability to support guideline-consistent multi-step reasoning. In this paper, we propose an agentic AI framework for generating evidence-based second medical opinions by integrating Retrieval-Augmented Generation (RAG) with a controller-driven orchestration architecture. The framework indexes clinical guidelines through overlapping document chunking and vector embeddings, retrieves pertinent guideline evidence at inference time through similarity search, and coordinates multiple large language models (LLMs), including GPT-4o, Claude Opus 4.6, and Gemini 2.5 Flash, to produce structured recommendations aligned with clinical guidelines. A web-based interface allows users to submit structured patient information and obtain standardized outputs encompassing condition status, diagnosis, severity assessment, and treatment planning. Evaluation on representative simulated respiratory scenarios shows that the agentic pipeline achieves strong reasoning-level guideline agreement under a conservative G-Eval threshold ( τ = 0.90 ), matching the strongest tested single-model configuration (precision = 1.00, recall = 0.75, specificity = 1.00, F1-score = 0.86), while maintaining an average token-based cost of USD 0.0473 per case across four benchmark cases, mean end-to-end latency of 44.66 s in the tested environment, strong semantic alignment with guideline-derived reference answers (mean cosine alignment = 0.7629), and stable repeatability across 30 controlled executions (coefficient of variation (CV) = 2.8%). These findings suggest that, in the evaluated respiratory setting, controller-based multi-LLM orchestration can improve guideline agreement and output consistency relative to the tested single-model configurations.