MediRAG: Secure Question Answering for Healthcare Data
Emily Jiang, Alice P. Chen, Irene Tenison, Lalana Kagal · 2024
Retrieval augmented generation (RAG) allows large language models to answer domain-specific questions by using external knowledge bases without training on this domain data or fine-tuning on its updates. This is especially promising for clinical tasks, as medical data tends to be dynamic, private, and distributed. By exposing the source documents that inform the response, RAG enables greater interpretability as well as reduced hallucination, both of which are crucial for safe deployment in healthcare. However, applying standard RAG to answer questions across patient data is complicated by strict access control to sensitive records enforced by privacy regulations such as HIPAA as well as the distributed nature of EHRs among institutions. We propose MediRAG, a clinical QA system that (i) supports a hierarchical design for federated document retrieval, and (ii) enables policy-based access control (PBAC) at retrieval time. For our experiments, we use the MIMIC-IV dataset, a publicly available EHR database that has been used in a wide array of research studies. We carry out a simulation of multiple federated hospitals and show that our scheme results in no loss of quality against a centralized baseline. We also evaluate performance with respect to key RAG metrics such as ROUGE, BLUE, context relevance, faithfulness and answer relevance and show that MediRAG is effective at clinical question answering across decentralized EHR documents while enforcing policies on sensitive data.