ViTCoin: A Vision Transformer-based System for Bangladeshi Coin Detection to Assist the Visually Impaired and Numismatic Heritage
Partha Bhowmik, Niloy Bhowmik, Mohammad Shahidur Rahman · 2024
In this modern era, accurate and effective coin recognition to conduct autonomous financial transactions is very crucial for the visually impaired. In our paper, we conducted a deep-learning based Vision Transformer (ViT) coin detection system to recognize Bangladeshi Coins of 1 Poisha, 5 Poisha, 10 Poisha, 25 Poisha, 50 Poisha, 1 Taka, 2 Taka and 5 Taka denominations. Our system is trained on a dataset of over 14,000 images in which coins are captured in a variety of lighting, orientation, and backdrop conditions. Notably, the application of any state-of-the-art Vision Transformer (ViT) In the realm of Bangladeshi Coin recognition has not been explored by any prior researchers. Moreover, the lackings and limitations of previous coin datasets, face challenges as training any models, which leads to poor performance in real-world and noisy images. In addition to making practical identification easier, the method also highlights how crucial it is to acknowledge and preserve the cultural and historical significance of Bangladeshi numismatic heritage. The proposed system detects both sides of the coins. It achieves a high level of accuracy, with test results indicating up to a remarkable maximum test accuracy of 99.39% in determining the proper denomination. This accuracy ensures reliable use in everyday settings where visually impaired people frequently struggle to identify currencies.