An Improved Masking Strategy for Self- Supervised Masked Reconstruction in Human Activity Recognition
Jinqiang Wang, Wenxuan Cui, Tao Zhu, Huansheng Ning, Zhenyu Liu · IEEE Sensors Journal · 2024
Masked reconstruction serves as a crucial pretext task for self-supervised learning, enabling the model to improve its feature extraction capabilities by reconstructing the masked segments from abundant unlabeled data. In the field of human activity recognition, this pretext task has traditionally employed a masking strategy that focuses on the time dimension. However, this strategy fails to fully leverage the inherent characteristics of wearable sensor data and overlooks the inter-channel information coupling, thereby limiting its potential as an effective pretext task. To overcome these limitations, we propose a novel masking strategy called Channel Masking. This strategy involves masking the sensor data along the channel dimension, thereby encouraging the encoder to extract channel-specific features while performing the masked reconstruction task. Additionally, Channel Masking can be seamlessly integrated with masking strategies along the time dimension, motivating the self-supervised model to undertake the masked reconstruction task in both the time and channel dimensions. The integrated masking strategies are named Time-Channel Masking and Span-Channel Masking. Furthermore, we optimize the reconstruction loss function to incorporate the reconstruction loss in both the time and channel dimensions. We evaluate the proposed masking strategies on three public datasets, and the experimental results demonstrate that these strategies outperform prior approaches in both self-supervised and semi-supervised scenarios.