Robust and Real-Time Bangladeshi Currency Recognition: A Dual-Stream MobileNet and EfficientNet Approach
Subreena, Mohammad Amzad Hossain, Mirza Raquib, Saydul Akbar Murad, Farida Siddiqi Prity, Muhammad Hanif, Nick Rahimi · IEEE Access · 2026
Accurate currency recognition is essential for assistive technologies, particularly for visually impaired individuals who rely on others to identify banknotes. This dependency puts them at risk of fraud and exploitation. To address these challenges, we first build a new Bangladeshi banknote dataset that includes both controlled and real-world acquiring scenarios. To further enhance the dataset’s robustness, we incorporate four additional datasets, including public benchmarks, to cover various complexities and improve the model’s generalization. To overcome the limitations of current recognition models, we propose a novel hybrid CNN architecture that combines MobileNetV3-Large and EfficientNetB0 for efficient feature extraction. This is followed by an effective multilayer perceptron (MLP) classifier to improve performance while keeping computational costs low, making the system suitable for resource-constrained devices. The experimental results show that the proposed model achieves 98.25% accuracy on controlled datasets, 93.33% on complex backgrounds, and 94.40% accuracy when combining all datasets. The model’s performance is thoroughly evaluated using five-fold cross-validation and seven metrics: accuracy, precision, recall, F1-score, Cohen’s Kappa, Matthews Correlation Coefficient (MCC), and Area Under the ROC Curve (AUC). Paired t-tests confirm statistical significance over seven baseline architectures in 32 of 35 comparisons ( $p \lt 0.05$ ), while an ablation study validates each pipeline component of the proposed system. Additionally, explainable AI methods like LIME and SHAP are incorporated to enhance transparency and interpretability. The system is deployed as a voice-enabled web application that provides predictions in English and Bangla, with end-to-end inference latency of 716–760 ms on a CPU-only research instance and full end-to-end response time of 800–900 ms from image capture to voice output on the deployment laptop. The code for this research can be found in this link: https://github.com/subreena/bangladeshi