Enhancing Medical Summarization with Parameter Efficient Fine Tuning on Local CPUs

Shamus Sim Zi Yang, Goh Man Fye, Wei Chung Yap, D. Yu · 2024

Documenting and summarizing patient symptoms and medical history for each visit can significantly burden clinicians' time management. Large Language Models (LLMs) have demonstrated great potential in natural language processing (NLP) tasks; however, their effectiveness in clinical summarization tasks has not yet been rigorously validated. While much research has focused on leveraging closed LLMs like GPT-4, Claude, and Gemini for clinical applications, privacy concerns hinder their deployment in real clinical settings. On-premises deployment offers a potential solution. This study examines domain adaptation techniques on the open-source LLM, Llama 3 8B Instruct, to improve clinical summarization. Our approach emphasizes fine-tuning on CPUs instead of the more commonly used GPU s, aiming for greater cost savings in practical applications. We apply Quantized Low-Rank Adaptation (QLoRA) for efficient task-specific adaptation and introduce CPU optimization techniques such as IPEX-LLM and Intel® AMX to enhance performance. Our results show that CPU fine-tuning provides a practical, cost-effective, and privacy-aware alternative to GPU fine-tuning, while improving the accuracy of medical summarization and enabling customization to meet unique clinical requirements.

Read the paper · More papers on PaperTik