Multimodal Human Activity Recognition Using Contrastive Fusion Learning and Lightweight Isomorphic Encoder for IoT-Enabled Smart Homes
Qi-Sen Hong, Ching-Hu Lu · IEEE Internet of Things Journal · 2025
Human activity recognition (HAR), integrated with the Artificial Intelligence of Things (AIoT) technologies, enables real-time user behavior detection across a wide range of applications. However, existing multimodal HAR approaches are costly, requiring a substantial amount of labeled data and expert annotations. The baseline methods have utilized contrastive learning yet they fail to account for negative sample misclassification, and they design separate encoders per modality without considering the architectural complexity. To address these challenges, we propose contrastive fusion learning with negative set pruning (CFL-NSP), which improves upon contrastive learning-based multimodal HAR models by effectively handling misclassified negative samples. Additionally, we introduce a lightweight multi-modal feature isomorphic encoder (Li-MFIE), which unifies multimodal sensor data into a common image format, achieving improved performance over models that require separate encoders for each modality. Our experiments utilize accelerometers, gyroscope, and skeleton-based data to provide a thorough multimodal HAR analysis. Experimental results show that combining both techniques improves accuracy from 53.88% to 58.86% (+4.98%) with a 5% label rate. The CFL-NSP improves accuracy from 53.64% to 57.74% (+4.10%) by removing false negatives. On the other hand, the Li-MFIE increases accuracy from 53.88% to 55.81% (+1.93%). Additionally, using the Li-MFIE reduces FLOPs from 3775.81M to 1519.79M (-59.53%) and decreases model training time from 105s to 84s (-20.00%). These results demonstrate its potential impact on applications such as smart homes, healthcare monitoring, and fitness tracking, where real-time multimodal HAR can enhance safety and user experience. With the unified input data format, our approach has the potential to be scalable and generalizable to various HAR scenarios without additional encoder design, making it suitable for a wider range of AIoT applications.