To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question Answering
Giacomo Frisoni, Alessio Cocchieri, Alex Presepi, Gianluca Moro, Zaiqiao Meng · 2024
Medical open-domain question answering demands substantial access to specialized knowledge.Recent efforts have sought to decouple knowledge from model parameters, counteracting architectural scaling and allowing for training on common low-resource hardware.The retrieve-then-read paradigm has become ubiquitous, with model predictions grounded on relevant knowledge pieces from external repositories such as PubMed, textbooks, and UMLS.An alternative path, still under-explored but made possible by the advent of domain-specific large language models, entails constructing artificial contexts through prompting.As a result, "to generate or to retrieve" is the modern equivalent of Hamlet's dilemma.This paper presents MEDGENIE, the first generate-thenread framework for multiple-choice question answering in medicine.We conduct extensive experiments on MedQA-USMLE, MedMCQA, and MMLU, incorporating a practical perspective by assuming a maximum of 24GB VRAM.MEDGENIE sets a new state-of-the-art in the open-book setting of each testbed, allowing a small-scale reader to outcompete zero-shot closed-book 175B baselines while using up to 706× fewer parameters.Our findings reveal that generated passages are more effective than retrieved ones in attaining higher accuracy.1 * Equal contribution (co-first authorship). 1 Our code, fine-tuned models, and generated contexts are publicly available at https://github.com/unibo-nlp/medgenie.