Large Language Model Empowered Privacy-Protected Framework for PHI Annotation in Clinical Notes

Guanchen Wu, Linzhi Zheng, Han Xie, Zhen James Xiang, Jiaying Lu, Darren Liu, Delgersuren Bold, Bo Li, Xiao Hu, Carl Yang · Studies in health technology and informatics · 2025

De-identifying private information in medical records is crucial to prevent confidentiality breaches. Rule-based and learning-based methods struggle with generalizability and require large annotated datasets, while LLMs offer better language comprehension but face privacy risks and high computational costs. We propose LPPA, an LLM-empowered Privacy-protected PHI Annotation framework, which uses few-shot learning with pre-trained LLMs to generate synthetic clinical notes, reducing the need for extensive datasets. By fine-tuning LLMs locally with synthetic notes, LPPA ensures strong privacy protection and high PHI annotation accuracy. Experiments confirm its effectiveness, efficiency, and scalability.

Read the paper · More papers on PaperTik