FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition
Yuxuan Weng, Tianyue Zheng, Yanbing Yang, Jun Luo · IEEE Transactions on Mobile Computing · 2025
Radio-Frequency (RF)-based Human Activity Recognition (HAR) rises as a promising solution when low-light, obstructions, or privacy concerns render computer vision impractical. However, thescarcityof labeled RF data due to their non-interpretable nature poses a significant obstacle. Thanks to the recent breakthrough offoundation models (FMs), extracting deep semantic insights from unlabeled visual data become viable, yet these vision-based FMs fall short when applied to small RF datasets. To bridge this gap, we introduce FM-Fi 2.0, an innovative cross-modal framework engineered to translate the knowledge of vision-based FMs for enhancing RF-based, multi-person HAR systems. FM-Fi 2.0 first employs the intrinsic capabilities of FM and RF modality to associate both intra- and cross-modal features of each subject, while simultaneously filtering out irrelevant features to achieve better alignment between the two modalities. FM-Fi 2.0 also employs a cross-modalcontrastiveknowledge distillation mechanism, enabling an RF encoder to inherit the interpretative power of FMs for achieving zero-shot learning. The framework is further refined through metric-based few-shot learning techniques, aiming to boost the performance for predefined HAR tasks. Comprehensive evaluations evidently indicate that FM-Fi 2.0 rivals the effectiveness of vision-based methodologies, and the evaluation results provide empirical validation of FM-Fi 2.0's generalizability across various environments.