An Empirical Study of Named Entity Recognition for Electronic Health Records

Nichapha Manoonwong, Anca Muresan, Anudeep Reddy Raavi, Xingquan Zhu · 2024

Named Entity Recognition (NER), which locates and recognizes phrases (i.e., entities) carrying specific meaning, such as locations, organization names, etc., is an important information extraction step for domains such as the processing of biomedical or Electronic Health Records (EHRs). To date, many methods exist for NER, and neural language models have emerged as the most sought-after tool due to their superb performance compared to other alternatives, such as rule-based approaches. Nevertheless, the rich set of tools and algorithms also raises concerns for both researchers and practitioners: how algorithms vary in their performance and what the general practices are in selecting proper NER algorithms for EHR processing. In this paper, we conduct an empirical study to understand the performance of different types of neural language models for EHR named entity recognition. Our main goal is to infer the performance of each model type with respect to the sample volumes and distributions with a special emphasis on context-dependent and low-frequency entities that pose a significant challenge, even for state-of-the-art models. Five types of models, LSTM, BiLSTM, Basic BiLSTM-CRF, Enhanced BiLSTM-CRF, and BERT, are studied in our experiments by using the N2C2 dataset (unstructured notes from the research patient data registry) as the test-bed. We vary the sample volumes and distributions and comparatively study the model performance. Our study draws important findings for researchers to decide the most suitable NER tools for EHRs.

Read the paper · More papers on PaperTik