DoctorAgent-RL: A Multi-Agent Collaborative Reinforcement Learning System for Multi-Turn Clinical Dialogue
Yichun Feng, Jiawei Wang, Lu Zhou, Zhen Lei, Yixue Li · 2026
Large language models (LLMs) excel at biomedical question answering but struggle in real clinical consultations. Single-round systems require patients to list all symptoms initially, often causing vague diagnoses. Traditional multi-turn models, limited by static supervised learning, lack flexibility and cannot intelligently gather key clinical data. To overcome this, we propose DoctorAgent-RL, a multi-agent reinforcement learning (RL) framework that treats medical consultations as dynamic decision-making under uncertainty. The doctor agent optimizes its questioning strategy via multi-turn interactions with a patient agent, dynamically adjusting its information collection based on rewards from a Consultation Evaluator. This RL fine-tuning allows LLMs to develop clinical reasoning strategies, not just mimic existing dialogues. We also built MTMedDialog, a new English multi-turn medical dataset designed for interactive simulation. It contains detailed case profiles that allow a patient agent to progressively reveal symptoms in response to the doctor’s questions, offering a more realistic training and testing setting than static datasets. Experiments show DoctorAgent-RL outperforms existing models in diagnostic accuracy. This approach reduces the risk of misdiagnosis in time-sensitive situations, frees clinicians to focus on complex cases, and helps optimize the use of medical resources.