LactFormer: Language-Aware Context Learning With Multi-Cue Transformers for HOI Detection

Weilong Peng, Qingfeng Chen, Keke Tang, Zhaoyang Yu, Yangtao Wang, Yanzhao Xie, Meie Fang · IEEE Transactions on Consumer Electronics · 2025

Human-Object Interaction (HOI) detection is critical for advancing machine understanding of complex visual environments involving human activities and their interactions with objects. Existing Transformer-based approaches have achieved significant progress but often fall short in fully leveraging prior knowledge and semantic context. To address these challenges, we propose language-aware context learning with multi-cue Transformers (LactFormer), a novel two-stage framework that enhances HOI detection by integrating multi-granularity cues from both visual and linguistic domains. LactFormer comprises three key components: the Language-aware Context Learning Module (LCLM), Vision-based Interactive Query Generation (VIQG), and the Language-aware Interaction Embedding Module (LIEM). The LCLM encodes positional language descriptions and merges them with visual context via a Transformer Encoder, forming a rich, language-aware context. This language-aware context provides a robust foundation for the queries obtained by VIQG, which are then processed by the Transformer Decoder in LIEM to generate effective interaction embeddings for human-object pairs. LactFormer overcomes the limitations of previous Transformer-based models by leveraging a comprehensive understanding of both visual and linguistic information. Extensive experiments demonstrate that LactFormer significantly improves interaction recognition accuracy, achieving state-of-the-art performance.

Read the paper · More papers on PaperTik