Study of Pre-trained Language Models for Named Entity Recognition in Clinical Trial Eligibility Criteria from Multiple Corpora

Jianfu Li, Qiang Wei, Omid Amir Ghiasvand, Miao Chen, Victor S. Lobanov, Chunhua Weng, Hua Xu · 2021

Named entity recognition (NER) is fundamental to the computerization of clinical trial eligibility criteria text. In this study, we fine-tuned pretrained contextual language models to support the NER task on clinical trial eligibility criteria. We systematically explored four pre-trained contextual embedding models for biomedical domain (i.e., BioBERT, BlueBERT, PubMedBERT, and SciBERT) and two models for the open domain (BERT and SpanBERT), for NER tasks using three existing clinical trial eligibility criteria corpora. Our evaluation results showed domain-specific transformer models achieved better performance than the general transformer models, with the best performance obtained by the PubMedBERT model (F1-scores of 0.715, 0.836, and 0.622 for the three corpora respectively. This study demonstrated the efficiency of domain-specific transformer-based language models for NER in clinical trial eligibility criteria.

Read the paper · More papers on PaperTik