Arkangel AI: A conversational agent for real-time, evidence-based medical question-answering

Maria Camila Villa, Natalia Castano-Villegas, Isabella Llano, Julian Martinez, Maria Fernanda Guevara, Jose Zea, Laura Velásquez · Intelligence-Based Medicine · 2025

ABSTRACT Introduction Large Language Models (LLMs) have been trained and tested on several medical question-answering (QA) datasets built from medical licensing exams and natural interactions between doctors and patients to fine-tune them for specific health-related tasks Objective We aimed to develop LLM-powered Conversational Agents (CAs) equipped to produce fast, accurate, and real-time responses to medical queries in different clinical and scientific scenarios. This paper presents MedSearch, our first conversational agent and research assistant. Methods The model is based on a system containing five LLMs; each is classified within a specific workflow with pre-defined instructions to produce the best search strategy and provide evidence-based answers. We assessed accuracy, intra/inter-class variability, and Cohen’s Kappa using the question-answer (QA) dataset MedQA. Additionally, we used the PubMedQA dataset and assessed both databases using the RAGAS framework, including Context, Response Relevance, and Faithfulness. Traditional statistical analysis was performed with hypothesis tests and 95% IC. Results Accuracy for MedQA (n: 1273) was 90.26% and Cohen’s kappa was 87%, surpassing current SoTAs for other LLMs (GPT-4o, MedPaLM2). The model retrieved 80% of the expected articles and provided relevant answers in 82% of PubMedQA. Conclusion MedSearch showed proficient retrieval and reasoning abilities and unbiased responses. Evenly distributed medical QA datasets to train improved LLMs and external validation for the model with real-world physicians in clinical scenarios are needed. Clinical decision-making remains in the hands of trained healthcare professionals.

Read the paper · More papers on PaperTik