Development of Language Model on Biomedical Domain to Pretrain Natural Language Processing

Vijaya Gunturu, Yadavalli Devi Priya, Gayatri Vijayendra Bachhav, K. Praveena, R J Anandhi, Navdeep Dhaliwal · 2024

Large neural language model like BERT can be pre trained to get extraordinary profits through multiple natural language processing task. Though, General Domain Corpora including web and news wire are focused on pre training efforts. The main specific pre training are benefited from general domain language models is considered as a prevailing assumption. The study focusses on the domain specific language model with abundance of unlabeled text like biomedical natural language processing and pre training from its scratch that results in more gains over the general domain language model. The investigation can be facilitated by compiling of biomedical NLP data sets that are publicly available. The experiment shows the pre training of domain specific model that act as a solid foundation in performing biomedical NLP task in wide range. the model is evaluated for modelling choices including task specific fine tuning and pre training. BERT models have some common practises involving named entity recognition using complex tagging schemes. The research can be accelerated with biomedical NLP for pre training and task specific model for the biomedical community and the leader board is created for biomedical language understanding and reasoning benchmark (BLURB).

Read the paper · More papers on PaperTik