Vision-Language Feature Refinement for Zero-Shot Object Counting

Md Jibanul Haque Jiban, Abhijit Mahalanobis, Niels da Vitoria Lobo · 2024

Large-scale vision-language pretraining has revolutionized zero-shot learning tasks by enabling models to effectively handle novel object categories using rich, generalized visual-language associations. However, it is challenging to use for dense tasks like zero-shot object counting due to their inherent difficulty in accurately localizing and quantifying a large number of unseen and diverse objects in complex images. To address this, we utilize a cross-modal encoder to learn joint representations from both modalities at the pixel level. We also introduce a refinement process that further processes the visual features to improve localization and capture multi-scale contextual information for robust counting. We perform extensive experiments on FSC-147 to show state-of-the-art performance on zero-shot counting and demonstrate improved cross-data generalization on the CARPK dataset.

Read the paper · More papers on PaperTik