Comparative Analysis of Fine-tuned Deep Learning Language Models for ICD-10 classification task for Bulgarian Language
Boris Velichkov, Sylvia Vassileva, Simeon Gerginov, Boris Kraychev, Ivaylo N. Ivanov, Philip Ivanov, Ivan Koychev, Svetla Boytcheva · 2021
The task of automatic diagnosis encoding into standard medical classifications and ontologies is of great importance in medicine -both to support the daily tasks of physicians in the preparation and reporting of clinical documentation, and for automatic processing of clinical reports.In this paper, we investigate the application and performance of different deep learning transformers for automatic encoding in ICD-10 of clinical texts in Bulgarian.The comparative analysis attempts to find which approach is more efficient to be used for finetuning of pre-trained BERT family transformer to deal with a specific domain terminology on a rare language such as Bulgarian.On the one hand, we use SlavicBERT and Mul-tiligualBERT models, which are pre-trained for a common vocabulary in Bulgarian but lack medical terminology.On the other hand, we compare them to BioBERT, ClinicalBERT, SapBERT, BlueBERT models, which are pretrained for medical terminology in English, but lack training for language models in Bulgarian, and vocabulary in Cyrillic.In our research study, all BERT models are fine-tuned with additional medical texts in Bulgarian and then applied to the classification task for encoding medical diagnoses in Bulgarian into ICD-10 codes.A big corpus of diagnoses in Bulgarian annotated with ICD-10 codes is used for the classification task.Such an analysis gives a good idea of which of the models would be suitable for tasks of a similar type and domain.The experiments and evaluation results show that both approaches have comparable accuracy.