Large Language Model Locally Fine-tuning (LLMLF) on Chinese Medical Imaging Reports
Junwen Liu, Zheyu Zhang, Jifeng Xiao, Zhijia Jin, Xuekun Zhang, Yuanyuan Ma, Fuhua Yan, Ning Winston Wen · 2023
The emergence of large language models exerts significant impact in the field of natural language processing. These models, which are based on attention networks, have shown remarkable capabilities in understanding and generating human conversation context, surpassing most of the state-of-the-art models in Natural Language Processing (NLP). Though Generative Pre-trained GPT-series open their public Application Programming Interfaces (APIs) for direct inference, such services require users to upload (relinquish) their data to servers, thus not suitable for domains operating sensitive data, such as the medical field. Due to comparable smaller model size, existing local pre-trained model deployments are usually not that helpful on domain-specific data without further fine-tuning. In this paper, we leverage Microsoft’s recent open source DeepSpeed-Chat platform and one of our selected pre-trained models, to locally conduct fine-tuning on our medical imaging reports (in Chinese), to infer diagnosis recommendations based on the imaging description. Our output models show very promising results according to both NLP metrics and experts’ evaluation criteria. This paper provides an initial report on our latest progress on locally fine-tuning large language model for medical data.