Enhanced Biomedical Named Entity Recognition Using SpaCy and BERT Models
Manish Kumar, Pardeep Singh, Poonam Kashtriya · Procedia Computer Science · 2025
Identification of named entity is a crucial procedure which entails locating and categorizing specified items in text, including people, places, organizations, medical codes, time expressions, amounts, and monetary amount. This paper concentrates on biomedical NER, which is vital for accurately identifying entities like diseases, chemicals, and drugs within biomedical research and clinical records. NER is crucial for applications including question-answering systems and automated extraction of information from large text corpora, as manually pinpointing these entities is both laborious and difficult. Automating this task with NER improves both efficiency and precision. The study evaluates and compares three biomedical NER models: SciBERT, a BERT-based model designed for scientific terminology; BlueBERT, trained with MIMIC-III clinical records and PubMed papers and SpaCy, a model pretrained with biomedical data. The methodology involves preprocessing the BioCreative V CDR (BC5CDR) biomedical text corpus for extracting chemical-disease relationships, annotating entities such as chemicals and diseases, getting the training dataset ready, and running the models through several training epochs. Performance is measured using metrics like accuracy, precision, recall, and F1-score, which provide insights into the effectiveness of each model architecture in biomedical NER. Results demonstrate that SciBERT having a remarkable 96.81% F1-score, followed by BlueBERT at 95.75%, and SpaCy at 90.96%. These results offer valuable guidance for researchers in choosing the most suitable model for developing automated systems to identify certain entities in the medical domain.