Physically plausible data augmentations in contrastive learning for wearable IMU-based activity recognition
Nobuyuki Oishi, Philip M. Birch, Paula Lago, Daniel Roggen · Knowledge-Based Systems · 2026
Wearable inertial measurement unit (IMU)-based human activity recognition (HAR) is a key enabler of mobile and pervasive computing applications. To address the scarcity of labeled data, self-supervised contrastive learning has been adopted. It leverages data augmentations to generate diverse views of sensor data, using them as positive pairs in a pretext task to learn representations from unlabeled data. However, conventional signal transformation-based data augmentations (STDAs) are poorly grounded in the realistic variability that affects IMU readings, e.g., diverse motion patterns and device positioning, often resulting in unstable downstream task performance. In this work, we investigate the effectiveness of ensuring physical plausibility in data augmentations within contrastive learning. Specifically, we compare physically plausible data augmentations (PPDAs), realized via physics simulation to incorporate domain knowledge, against legacy STDAs. Evaluated on three public datasets covering daily activities and fitness workouts (8 to 34 classes), models pretrained with PPDAs consistently outperform STDA-pretrained models and fully supervised baselines, improving the downstream classification macro-F1 score by 3 to 14 percentage points. Qualitative analyses reveal that PPDAs shape the learned representation space into a more physically grounded topology. Furthermore, PPDAs achieve fully supervised-level performance with less than 10% of labeled data. We also demonstrate the robustness of physically grounded augmentations, showing they remain effective with lower-fidelity simulation data and generalize across different encoder architectures and contrastive learning frameworks. These findings suggest that incorporating physical domain knowledge to ensure physical plausibility enables the generation of semantically consistent positive pairs, helping models learn representations more relevant to HAR.