Classifying Bangla Book’s Context: A Multi-Label Approach

Md Safaet Ullah, Md. Ali Al - Reza, Mohammad Shahidur Rahman · 2023

This paper thoroughly analyses the multi-label classification of Bangla book content. We introduce a novel dataset of the Bangla book’s context, annotated with multiple labels, based on eight genres. Bangla is one of the most widely spoken languages and has an extensive literary tradition. However, the classification of book context data remained largely unexplored due to insufficient pertinent data. To overcome this limitation, we present Bangla Book Context, a large-scale dataset containing 10,211 samples categorized into eight main categories: fiction, nonfiction, mystery, tragedy, thriller, motivational, and romantic. For a proper evaluation, we employed three pre-trained models: bert-base-multilingual-case, BanglaBERT, BERT: Pre-training of Deep Bidirectional Transformers for Language Model and CNN-LSTM attention-based deep learning approach. However, the CNN-LSTM attention-based deep learning model consistently performs better than all the other models we have described.

Read the paper · More papers on PaperTik