An Agentic LLM Framework for Autonomous Surgical Continuum Monitoring: ReAct-Driven Tool-Use Agents for Presurgical, Intraoperative, and Postsurgical Cardiopulmonary Care
Charalampia Pylarinou, Lefteris Gortzis, Vasileios Leivaditis, Elias Liolis, Andreas Antzoulas, Spyros Papadoulas, Konstantinos Nikolakopoulos, Ioannis Panagiotopoulos, Sofoklis Mitsos, Periklis Tomos, Efstratios N. Koletsis, Francesk Mulita · Bioengineering · 2026
BACKGROUND: Rule-based multi-agent system (MAS) architectures for healthcare coordination rely on hardcoded decision trees that cannot generalise to novel clinical scenarios or self-correct reasoning errors. These limitations are acute in surgical continuum care, where patients traverse presurgical risk stratification, intraoperative monitoring, postsurgical ICU, ward care, and remote rehabilitation over days to weeks-a complexity no fixed-policy agent architecture can address without prohibitive rule engineering. OBJECTIVE: We present the first agentic large language model (LLM) framework for autonomous end-to-end surgical continuum monitoring, superseding the prior rule-based MAS Digital Twin. Six ReAct-driven tool-use agents replace fixed-policy agents with dynamic reasoning, multi-hop evidence retrieval, and Reflexion self-correction while maintaining mandatory confidence-gated Human-in-the-Loop (HITL) gating at every care-pathway-modifying decision. METHODS: The framework is grounded in the ReAct paradigm and Reflexion self-evaluation, embedded within the DETER Digital Twin state engine S(t). Each agent is specified by a ReAct loop signature, a ten-function clinical tool registry, and confidence-gated HITL escalation logic. Inter-agent coordination replaces the rule-based Priority Queue Manager with an LLM-mediated Coordination Supervisor Agent reasoning over competing resource requests. RESULTS: The framework delivers: (i) six formally specified ReAct-loop agents with explicit tool registries and authorisation boundaries; (ii) a confidence-gated HITL architecture that reduces alert fatigue while preserving safety for ambiguous clinical scenarios; (iii) an extended conflict resolution function P(p,t,context) incorporating surgical phase and DETER deterioration trajectory gradient; (iv) Reflexion self-correction with a formal N_max = 2 termination condition and Clinical Factuality Verification Layer; and (v) a multi-phase Digital Twin state engine extending S(t) to the full surgical continuum. CONCLUSIONS: The proposed framework represents a fundamental architectural departure from rule-based clinical AI-from hardcoded policies to dynamic reasoning, from static retrieval to multi-hop tool-use chains, and from fixed escalation thresholds to confidence-gated self-evaluation-providing a formally specified, clinically deployable foundation for next-generation autonomous surgical care coordination.