Span-based annotation framework for LLM-based clinical named entity recognition: development and validation using Korean emergency department notes
Eun Hye Jang, Javier Aguirre, Sangji Lee, Hyeyoon Moon, Won Chul Cha · JAMIA Open · 2025
Objective: This study aims to develop and validate of a span-based annotation framework for clinical named entity recognition (NER) using large language models (LLMs) based on Korean emergency department clinical notes. Materials and Methods: Two datasets with the same entity types but different annotation spans (word- vs phrase-level) were constructed, with the phrase-level dataset further was expanded into a doubled version. A Korean language-specific LLM was fine-tuned on each dataset, producing three variants that were compared with two baseline models, few-shot LLM and fine-tuned small language model (SLM). The final variant fine-tuned on the doubled phrase-level dataset was further evaluated against a human annotator. Results: In all experimental settings, three variants outperformed the baselines by achieving the highest F1 scores across all metrics. The final variant achieved F1 scores exceeding 0.80 across all averaging strategies and evaluation metrics, including token-based, span-based exact, and span-based partial evaluations demonstrating its robustness applicable in a practical setting. Discussion: While prompt engineering with few-shot is widely adopted for LLM-based clinical NER, our results proved that supervised fine-tuning (SFT) is consistently superior. The final variant outperformed the human annotator, emphasizing its potential as an automatic labeling tool. Conclusion: This study introduced a novel span-based annotation framework for LLM-based clinical NER verified by three independent experiments. In multilingual and real-world clinical settings, LLMs have proven in handling complex entity spans that include word-level and phrase-level annotations, particularly for long and attribute-rich entities.