Enhancing the capabilities of Llama2 in biomedical applications
ren yuankai, Xie Zhenping · 2024
This study is dedicated to the development and evaluation of a large language model specifically optimized for the biomedical domain, using the sophisticated Llama2 architecture as its foundation. By delving into the vast corpus of biomedical literature, which includes medical research papers, news reports, drug labels, and case studies, we have curated a comprehensive training dataset comprising 10 billion tokens aimed at addressing the deficiencies of existing language models in specialized medical knowledge. During the model training phase, we initially conducted pre-training using data from the biomedical domain, followed by supervised training of a biomedical question-answering model utilizing LoRA technology. In the evaluation phase, the model's performance underwent rigorous testing with challenging datasets, including the Chinese National Medical Licensing Examination, the United States Pharmacist Licensing Examination, and the PubMedQA question-answering dataset. These tests served as a robust indicator of the model's practical applicability in real-world medical professional scenarios. Experimental results suggest that our model performs exceptionally well in various specific biomedical tasks and surpasses the well-known ChatGPT-3.5 model in key performance metrics, affirming the effectiveness of domain-optimized large language models. Despite the exemplary performance, the model still lags behind the latest ChatGPT-4, which indicates potential for improvement in understanding depth, reasoning capabilities, and knowledge updating. Overall, this research provides a potent tool for natural language processing tasks in the biomedical field and lays the groundwork for the future development of more specialized and precise large language models.