Enhancing Contextual Understanding in GPT through Multimodal Pre-training
Sonu Kumar, Simran, Nongmeikapam Thoiba Singh · 2023
Natural language processing (NLP) tasks that Generative Pre-trained Transformers (GPT) have proven to be extremely effective at include text production, machine translation, and question answering. These models mostly rely on textual data, despite the fact that there is a lot of information available in other modalities such as images, audio, and video. In order to enhance the contextual comprehension of GPT models, this study will look into the benefits of multimodal pretraining. These massive models simply receive training on simple texts without any linguistic or global comprehension, regardless of their achievement. Furthermore, the bulk of large-scale models were trained using auto-regression. Because of this, this traditional way of fine-tuning performs terribly on future language comprehension exams.