Enhancing zero-shot object counting via ViT and BERT integration

Penghang Lu, Ping Luo, Xiafu Lv · 2025

Computer vision-based counting methods effectively solve the measurement errors and high time consumption of traditional manual counting, reducing the workload of relevant personnel. However, most existing counting models are only applicable to specific categories and usually require manual examples when dealing with multi-category objects, which limits their application in automated systems. To this end, this paper proposes the ViTBERT-Count method, which uses a text-guided zero-shot counting system to enable the model to automatically count target categories based on text descriptions without labeled data. The model combines the Vision Transformer (ViT) and the Bidirectional Encoder Representation (BERT) to achieve efficient image and text feature alignment and information interaction capabilities. ViT extracts high-resolution image patch features, while BERT provides powerful language understanding and context modeling capabilities. This combination enables the model to perform well in zero-shot object counting tasks.

Read the paper · More papers on PaperTik