IMUZero: Zero-Shot Human Activity Recognition by Language-Based Cross Modality Fusion
Jie Su, Fengtong Ge, Zhenyu Wen, Taotao Li, Yang Bai, Yejian Zhou, Xiaoqin Zhang · Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies · 2025
Wearable-based human activity recognition (HAR) typically uses motion sensor data, such as inertial measurement unit (IMU) signals, to identify human movements. While effective in controlled scenarios, traditional HAR models are trained on a fixed set of activities and fail to generalize to new or unseen actions. This limitation motivates the use of zero-shot learning (ZSL), which aims to recognize unseen activities without direct training examples. Existing ZSL methods often rely on projecting seen and unseen classes into a shared latent space using external semantic information, such as visual or textual data. However, visual data are commonly unavailable in wearable settings, and text-based semantics from activity labels or coarse descriptions lack the detail needed for accurate recognition. Recent work explores large language models (LLMs) to provide prior knowledge through question-answering mechanisms. While promising, these approaches do not use raw sensor data directly and often miss important contextual signals. We propose IMUZero , a ZSL framework that fuses sensor signals with LLM-generated semantic attributes. Our method uses LLMs to produce fine-grained, decomposable activity attributes without additional LLM-based training, preserving sensor context. We also introduce a channel shuffle order constraint that models axial bias to improve generalization. Experiments on four public datasets show that our method outperforms existing ZSL approaches that rely on learned semantic embeddings. We release the code at https://github.com/Was-Lab/IMUZero.