LCNet-ViT-FG: a product recognition method based on the fusion of self-supervised pretrained CNN and transformer
Wenzhong Shen, Yinshen Qin · Journal of Electronic Imaging · 2025
Accurate product recognition is critical in real-world applications such as retail checkout and inventory management. However, high interclass similarity, large intraclass variation, and long-tailed data distributions pose significant challenges for existing methods, especially under limited annotation. We propose LCNet-ViT-FG, a self-supervised learning framework specifically optimized for fine-grained product recognition. Our model integrates convolutional and Transformer-based modules to combine local and global feature modeling and incorporates a self-supervised pretraining strategy to reduce reliance on large-scale labeled data. In addition, we introduce a hybrid loss function that combines CosFace and focal loss, enhancing class separability and improving the recognition of hard samples. Extensive experiments conducted on both public datasets and our real-world product dataset (SEPAD_RP) demonstrate that our method achieves superior accuracy and robustness, especially under occlusion, deformation, and class imbalance. We further investigate the model’s performance under noisy conditions, confirming its strong resilience to input noise. Multirun evaluations with different random seeds also validate the stability and reproducibility of our framework. The proposed LCNet-ViT-FG offers a promising solution for scalable and reliable product recognition in dynamic, real-world environments.