Optimization of Biomedical Language Model with Optuna and a Sentencepiece Tokenization for NER

Chérubin Mugisha, Incheon Paik · 2022 IEEE International Conference on Bioinformatics and Biomedicine (BIBM) · 2022

Raw medical documents come with in-domain challenges that need to be addressed with appropriate methods, such as a suitable tokenizer and tailored model via hyperparameters. Most available pre-trained models have been trained on clean biomedical data from published documents without considering the characteristics of raw medical texts. This approach drastically limits the performance of those models on real-world data such as electronic health record documents. This work introduces a fine-grained language model trained over a SentencePiece tokenizer with various biomedical and clinical data. Through this RoBERTa-based model, referred to as BioBERTa, we demonstrated the importance of data-dependant optimization by evaluating our model on several Named Entity Recognition tasks. Our approach showed state-of-the-art results on NCBI-disease and BC5-disease datasets. We demonstrated the effectiveness of an in-domain tokenizer where ours improved the embedding length by 15.1% of its original model. Our approach also shows that fine-tuning the hyperparameters can significantly improve the model accuracy(up to 2% F1 score).

Read the paper · More papers on PaperTik