GEN-CVCI: Optimization of the HOI Detection Model Based on GEN
Ke Yang · 2025
Human-Object Interaction (HOI) detection aims to identify people, objects, and their interactions in images, typically represented as. The proposed model, GEN-CVCI, enhances the baseline GEN-VLKT through three key improvements. First, to strengthen visual feature extraction, GEN-CVCI integrates the Convolutional Block Attention Module (CBAM) into the visual encoder, improving feature representation for humans and objects. Second, a new category query learning module is introduced before the interaction decoder to enhance logical reasoning. Third, for finer-grained interaction classification, the model combines outputs from the verb classifier and HOI classifier to refine predictions. On the V-COCO and HICO-DET benchmarks, GEN-CVCI achieves mAP improvements of 2.11 and 2.3 percentage points, respectively, over GEN-VLKT.