The open-world object counting algorithm based on YOLO-WOLD

xianyi mao · 2025

Counting is a frequently encountered problem, such as tallying the number of students who have arrived in the classroom, counting items on a shelf, or enumerating books on a bookshelf. In densely packed scenarios, manual counting can be quite cumbersome. However, if counting can be performed by a computer, efficiency can be significantly enhanced. Conventional detection models typically focus on specific categories of objects, lacking generalizability. These models require retraining for each new type, making cross-domain applications challenging. This paper introduces a counting algorithm named YWcount, based on the Yolo-World model. Since the foundational model has undergone pre-training on extensive datasets of image-text pairs, it has assimilated a wealth of semantic knowledge and visual characteristics. By freezing the image-text feature extraction module in Yolo-World, the algorithm leverages these pre-trained model features to effectively count new categories of objects without the need for extensive annotated data. Firstly, an image example extraction branch is added to the network structure, enabling the model to accept dual-modal input information to specify the objects to be counted, thereby enhancing the model's generalizability. Secondly, a cross-attention mechanism is employed to fuse text example and image example features, allowing the model to fully learn the intrinsic appearance information provided by visual examples. Finally, a density map regression counting loss head is introduced, which leverages the global object quantity information provided by the density map to enhance the detection accuracy of the model.

Read the paper · More papers on PaperTik