AI-Powered Retrieval-Augmented Generation Framework with Large Language Models for Enhanced Public Health Response

Shivam Bhardwaj, Rakesh Chandra Joshi, Vansh Tiwari, Kamil Říha, Pavel Sikora, Malay Kishore Dutta · 2025

Healthcare and biomedical industries produce massive amounts of data such as standard operating procedures, protocols, research papers, clinical guidelines, etc. The management and control of this amount of information can limit timely decisions in particular, during emergency situations (e.g., pandemics). Large Language Models (LLMs) can effectively manage this data, but naive implementations are often costly and time-consuming and have limitations such as hallucinations, outdated information, and irretractable thought processes. This research proposes a framework integrating LLMs with Retrieval-Augmented Generation (RAG) techniques and a vector database to enhance data retrieval and generation. The AI-based framework uses a variety of public health data sets consisting of structured and unstructured data, as well as full-text medical references. It addresses the challenges of implementing an RAG pipeline on medical data by developing open-source LLMs integrated with advanced retrieval methods. The framework achieves strong performance with ROUGE-1 precision of 0.6797, recall of 0.816, F-measure of 0.718, and BLEU of 0.4709. It generates accurate responses in 35 seconds, making it efficient for healthcare use. Tackling the key translational issues of deploying RAG pipelines to complex medical text, this study is generalizable to the wider use of LLMs in healthcare, disease prevention, and surveillance. Implementations of this system have the ability to change the world landscape of public health by providing epidemiologists with ongoing, up-to-date information to be used by healthcare practitioners, as well as by enhancing global disease surveillance and response capacity.

Read the paper · More papers on PaperTik