Study on the Efficacy of GPT Double Heads Model for Question-Answering in Bengali

Koustav Chakraborty, Suman Chowdhury, Rajesh P. Barnwal · 2024

Large Language Models (LLMs) represent the forefront of advancements in natural language processing, driving the development of numerous models based on either novel architectures or enhancements to existing ones. Among these, OpenAI’s GPT architectures stand out due to their flexibility and high customizability, making them ideal for specialization in diverse tasks. Despite ongoing progress, there remains a significant gap in developing robust NLP tools for low-resource languages like Bengali, which has over 250 million speakers. In this paper, we present a Bengali Question-Answering (QA) system utilizing the GPT Double Heads model, a transformer-based architecture. We curated and translated a dataset from the WikiQA corpus specifically for Bengali QA tasks and retrained the GPT Double Heads model on this dataset. Our evaluation demonstrates promising results, establishing a foundation for future research and applications in vernacular languages. By addressing the unique challenges posed by Bengali, this work contributes to the broader field of NLP by paving the way for more inclusive language technologies.

Read the paper · More papers on PaperTik