Language Model Embeddings to Improve Performance in Downstream Tasks
Abhishek Pawar, Parag Patil, Riya Hiwanj, Ajinkya Kshatriya, Diptee Vishwanath Chikmurge, Sunita Barve · 2024
Natural Language Processing (NLP) has made significant breakthroughs, mainly in the creation of transformer-based language models, and has demonstrated great performance in a variety of language comprehension tasks. Despite these achievements, maximizing the possibilities for downstream applications remains an ongoing problem. While models based on transformers excel at learning contextualized representations during pretraining, they can benefit from focused modifications in downstream tasks. Our study highlights the advantages of domain-specific model embeddings over fine-tuning language models based on generalized transformers. The study methodology entails a detailed examination of textual datasets, including Marathi news corpus, medical text, and legal references. Sophisticated cleaning procedures are applied to enhance model performance. The proposed end-to-end pipeline comprises validation, classification, RoBERTa fine-tuning, and data selection for downstream processes. Through a more nuanced knowledge of the evolving field of natural language processing (NLP), this research advances our understanding of the effectiveness and adaptability of refined transformer-based language model embeddings in real-world applications.