Context-Aware Open-Vocabulary HOI Detection Using Visual Prompt
Chunyan Chen, Zhu Teng, Yuanzhouhan Cao · 2025
Human-object interaction (HOI) detection plays an important role in scene understanding. Currently, HOI research faces challenges in recognizing novel interactions and handling complex, ambiguous scenarios. Traditional methods typically rely on predefined interaction categories, making them difficult to adapt to the dynamic changes and diversity of real-world scenes. While vision-language models have shown potential in bridging the gap between visual and textual features, existing methods still struggle to effectively utilize textual data, especially when dealing with complex interactions and occlusions. To address these challenges, a method combining context-aware semantic enhancement and a visual prompt strategy is proposed. In this approach, rich interaction descriptions are integrated with the pretrained CLIP model to achieve vision-text semantic alignment, effectively incorporating contextual information. Moreover, the visual prompt strategy dynamically aggregates visual features from interaction regions, optimizing interaction category embeddings and further improving the model's performance in complex scenarios. Experimental results on the HICO-DET and SWIG-HOI datasets demonstrate that this method outperforms existing approaches in terms of generalization, robustness, and handling unseen interaction categories.