A Deep Learning Approach for Multiclass Domain-Specific Question Classification in Computer Science for Bangla Language

Md Mosharaf Hossain, Dip Karmakar, Prapti Nandi, Md Osama, Md. Khalid Syfullah · 2025

Text classification is a pivotal challenge in Natural Language Processing (NLP), particularly for underrepresented languages like Bangla, spoken by over 230 million native speakers. Efficient categorization of Bangla text can facilitate rapid information retrieval and organization across diverse domains. However, the language’s unique structure and limited linguistic resources compared to English pose significant challenges. In this study, we introduce a novel dataset comprising domain-specific questions across various computer science fields, labeled with their respective domains. The dataset includes questions from 19 computer science courses such as Machine Learning, Digital Logic Design, and Data Structures, with over 20,000 samples in total. We benchmark several machine learning and deep learning models to classify these questions into their respective domains. Our findings demonstrate the effectiveness of these models in addressing domain-specific classification tasks, paving the way for enhanced information organization and domain-aware applications in Bangla text processing.

Read the paper · More papers on PaperTik