Semantic Context Re-Mining for Multimodal Guided Human-Object Interaction Detection

Jihao Dong, Hua Yang · 2025

Human-Object Interaction (HOI) detection aims to identify interactions between humans and objects in visual scenes. While vision-language models like CLIP have shown promising results in zero-shot HOI detection, challenges persist, including the limitation in spatial feature modeling and the semantic diversity of verbs in textual representations. To address these issues, we introduce a novel approach that utilizes multimodal prompts to guide interaction prediction. First, we re-mine the semantic context within visual features to generate a more comprehensive interaction representation. Second, we utilize pre-trained models to generate both visual and textual prompts, effectively transferring the prior knowledge to the HOI detection task. Our method achieves competitive performance on standard HOI benchmarks, Especially under the zero-shot setting, demonstrating its potential to advance the field of HOI detection.

Read the paper · More papers on PaperTik