Text-embeding and Visual-reference Less Object Counting

Yueying Tan, Linjiao Hou, Yutong Zhang, Guangyi Xiao, Jiaolong Wu · 2024

This study presents TET-Count, a novel category-agnostic model for object counting from natural language prompts, addressing limitations in existing methods requiring extensive annotated data. Leveraging pre-training of CLIP, TET-Count integrates feature projection and image regulation for enhanced semantic localization and counting accuracy. Experiments on the FSC-147 dataset using MAE and RMSE metrics show its superiority over other zero-shot techniques. Ablation studies confirm the importance of feature projection, and qualitative results demonstrate robust counting capabilities. TET-Count represents an advance in cross-modal object counting, combining NLP with deep learning for future research and applications.

Read the paper · More papers on PaperTik