BanglaMedNER: A Gold Standard Medical Named Entity Recognition Corpus for Bangla Text

Abdul Muntakim, Farhan Sadaf, K. M. Azharul Hasan · 2023

Medical Named Entity Recognition (MNER) has gained much attention recently and aims to recognize medical terms from text, such as proteins, diseases, symptoms, drugs, chemicals, and human anatomy. However, MNER for Bangla text has yet to progress significantly due to various challenges. One of these challenges is the need for a large and standardized corpus in this domain for Bangla. This work presents the annotated Bangla corpus for medical named entity recognition in standard IOB2 format consisting of more than 117000 tokens. The annotation process involved collecting texts from diverse Bangla health and drug-related domains. This corpus includes three distinct entity classes: Chemicals and Drugs (CD), Disease and Symptom (DS), and Anatomy (ANAT). Multiple machine-learning techniques were explored for medical named entity recognition, and the Bidirectional Long Short-Term Memory (BiLSTM) model, coupled with a Bidirectional Conditional Random Field (BI-CRF) layer, outperformed all the other models in our dataset with an outstanding F1 macro average score of 75%. We also employed class weights to account for the biases in our dataset.

Read the paper · More papers on PaperTik