An AI-based Approach to Train a Language Model on a Wide Range Corpus Using BERT
Aryan Bhutda, Gopal Sakarkar, Nilesh M Shelke, Ketan Paithankar, Riddhi Panchal · 2024
In 2018, Google developed and released BERT (Bidirectional Encoder Representation from Transformers) for Natural Language Processing (NLP). The English language models in BERT are two. To help close this data gap, a range of techniques have been developed for pre-training general purpose language representation models on the huge amount of unannotated content available on the web. Unlike when BERT is trained on these datasets from scratch, the model must first be pre-trained on small amounts of data. Once a measurable result is obtained, it can then be applied to tasks like NLP sentiment analysis and question answering, which causes to a notable rise in accuracy. The purpose of this research is to train the language model and conduct a descriptive study of BERT in various applications. For the implementation, throughout the task we are taking a famous textbook of medicine "Harrison’s Principle of Internal Medicine" into consideration. The research has made advantage of high configuration computation, specifically the Nvidia Tesla DGX-V100 GPU.