Shongket: A Comprehensive and Multipurpose Dataset for Bangla Sign Language Detection
Sk. Nahid Hasan, Md. Jahid Hasan, Kazi Saeed Alam · 2021 International Conference on Electronics, Communications and Information Technology (ICECIT) · 2021
Dataset is considered to be the most important resource in machine learning fields or computer vision-based applications. Image classification problems are heavily influenced by the quality of a dataset. Bangla Sign Language is a communication medium in Bangla that is used for communication between normal people and speech and hearing disabled people. As a result, recognition of Bangla Sign Language is vital for these disabled people. Any complete Bangla Sign Language dataset is extremely rare and fewer research works have been done in this field. To address this issue, we have created a complete dataset of Bangla sign language that includes all letters and digits, which we intend to make public for use in future research. We have captured 150 images per class for the 10 digit classes and 120 images per class for the 36 Bangla letter classes. In total, we have captured 5820 images for both digits and letters from volunteers. We have included as many variation as possible in our dataset to improve our training accuracy. We have tested the performance of various well-known deep learning models on our dataset which produced the maximum accuracy of around 95%. This dataset will help researchers in future Bangla sign language related works.