DPM-CLIP: Zero-Shot Multimodal Egocentric Activity Recognition based on Dual-Prediction Mechanism

Yukun Chen, Liang Wan, Zihuan Qiu, Mingzhou He, Fanman Meng, Linfeng Xu, Qingbo Wu, Hongliang Li · 2025

Advancements in Zero-shot Multimodal Egocentric Activity Recognition (ZS-MM-EAR) largely rely on Vision-Language Model (VLM). However, existing methods struggle with VLM’s inadequate representation of egocentric activities, including challenges in capturing egocentric-specific features, adapting to domain shifts between egocentric video and pre-training data, and effectively leveraging complementary data such as Inertial Measurement Unit (IMU). To address these issues, we propose DPM-CLIP, a ZS-MM-EAR method tailored for vision, text and IMU modalities. Firstly, we design an attribute-driven text augmentation module that leverages a Large Language Model (LLM) to generate fine-grained textual descriptions of activities. Secondly, we construct an Instance-feature Repository (IFR) to store base class features and generate pseudo-features for novel classes through feature center migration. Finally, we introduce a dual-prediction mechanism with a prediction correction module to enhance generalization and recognition accuracy. Extensive experiments on the UESTC-MMEA-CL dataset validate the effectiveness of the proposed method.

Read the paper · More papers on PaperTik