CORE-CLIP: Smart collaborative reasoning driven by CLIP for human-object interaction detection

Yuequan Yang, Haojun Zhang, Zhiqiang Cao, Jiarui Hong, Junzhi Yu, Xu Wang · Pattern Recognition · 2025

• We propose a smart semantic enhancement collaborative reasoning framework called CORE-CLIP, which is well driven by the openness and robustness of pre-training model CLIP. This framework effectively utilizes the semantic priors of visual-language pre-trained models to construct a text guided visual cross-modal alignment strategy to achieve fine-grained understanding of human-object interaction. The CORE-CLIP can achieve a good tradeoff between the model size and inference speed. • Two novel modules combining text features are elaborated for fine-grained interaction recognition and interaction classification. Specifically, the TDA module uses a dual fusion method to aggregate interactive text information and visual context in the structure to achieve accurate recognition of interaction categories, while the SEC module focuses on three special types of text information to explicitly and efficiently classify objects, actions, and interactions, respectively. • Extensive experiments demonstrate excellent performance on the two benchmarks and self-constructed validating set TableTennis. In particular, the proposed CORE-CLIP improves 1.36 mAP on V-COCO and 1.01 mAP on HICODET compared to GEN-VLKT. Zero-shot experimental results outperform the existing state-of-the-art methods such as GEN-VLKT and HOICLIP. The inference speed of CORE-CLIP surprisingly attains 38.3 FPS while its total parameters are only 52.9M. Ablation studies validate the contribution of each module for efficient and accurate interaction understanding. Human-object interaction (HOI) detection has attracted more and more attention due to its wide potential applications. Recently, Contrastive Language-Image Pre-training (CLIP) achieves promising results in 2D/3D zero-shot and few-shot learning. To overcome current heavy reliance on the large collection of annotated HOI data and often failure to recognize such unseen HOI relationships in training datasets in existing methods, we propose a smart collaborative reasoning framework named CORE-CLIP based on semantic enhancement driven by the openness and robustness of CLIP. Specifically, the text guided dual fusion attention module (TDA) captures precise interaction patterns by progressively integrating CLIP text embeddings, CLIP visual features, and global contextual features. The semantically enhanced explicit interaction classification module (SEC) is initialized by leveraging semantic information of objects, actions, and interactions generated through CLIP text embeddings, to ensure alignment between visual features and linguistic semantics. Extensive experiments demonstrate that CORE-CLIP has state-of-the-art performance on two benchmark datasets. In particular, compared to GEN-VLKT, the proposed CORE-CLIP improves 1.36 mAP on V-COCO and 1.01 mAP on HICO-DET. Zero-shot experimental results outperform state-of-the-art methods GEN-VLKT by a signifigant margin of 34.4 % in UO type setting and HOICLIP by 50.9 % in UO type setting, respectively. The inference speed of CORE-CLIP surprisingly attains 38.3 FPS while its total parameters are only 52.9M.

Read the paper · More papers on PaperTik