Medical Named Entity Recognition in Arabic Text using SVM
Rema Muftah Hamad, Ahmed Mohamed Abushaala · 2023
There is a huge amount of digitalized medical documents which are valuable for many applications. The text in a medical document is needed to collect knowledges and information about diseases and are used to make medical decisions. An automatic processing to extract information and recognizing named entities in these texts is very useful to facilitate many medical tasks. However, the text is not structured and is written in natural language, which makes this task a challenge. In fact, recognizing entities in Arabic language is not well studied in this field. In our work, we propose a Named Entity Recognition (NER) model to recognize entities such as: of disease, diagnosis, treatment, and symptom. For that, we perform a classification task using a dataset of 27 real medical documents based on a Support Vector Machine (SVM). We have pre-processed the text and selected candidate words to be classified. Then, we have represented each word by multiple features such as: FastText embedding, Part of Speech (PoS), Term Frequency-Inverse Document Frequency TF-IDF embedding, and non medical word matching. To classify the word, we take into account the surrounding words to capture its context. We have used PCA to reduce the size of TF-IDF vectors and we don’t use a hand-crafted supplemental resource. Thus, we have classified each word into its entity class by the SVM model. Evaluation results show that our model is stable and provides high and balanced results. In addition, we outperform a state-of-the-art model by an F1-score of +6.56% although the model requires a big manual effort for preparation.