Harformer: a channel-separated embedding unsupervised model for sensor-based human activity recognition

Zun Wang, Lingfei Mo, Yaojie Zhu · Measurement Science and Technology · 2025

Abstract Sensor-based human activity recognition (HAR) has garnered significant attention due to its wide range of applications, from healthcare to smart environments. Facing the challenge of difficult label collection in the HAR field, self-supervised learning has attracted significant attention due to its capability to extract features from data without relying on labels. Transformer-based models have achieved promising results in time-series tasks owing to their powerful performance. However, how to efficiently embed time-series data into tokens understandable by Transformer models remains a research-worthy problem. To reduce the parameters of the embedding layer in Transformer models and improve the feature extraction ability of the embedding layer for HAR data, in this paper, we propose Harformer, a novel unsupervised model that leverages a channel-separated mixed embedding (CSME) module and a patch-masking reconstruction strategy. The CSME module provides lightweight embeddings, significantly reducing computational complexity compared to traditional methods. By employing a reconstruction task as the unsupervised learning objective, Harformer effectively learns informative representations, achieving state-of-the-art performance on three public datasets: DSADS, PAMAP2, and MHEALTH. Experimental results demonstrate Harformer’s superiority over existing unsupervised models, with an average accuracy of 86.18% and an F1 score of 85.21%. Fine-tuning experiments further underscore the robustness of the pre-trained encoder, maintaining competitive performance with as little as 10% labelled data. Ablation studies validate the effectiveness of the CSME module, which outperforms convolutional neural network- and fully connected network-based embeddings while requiring fewer parameters. Harformer not only advances the state-of-the-art in unsupervised HAR but also lays a foundation for incorporating its components into diverse Transformer-based architectures. Future work will explore the generalisability of the CSME module across different Transformer encoder variants and time-series tasks, further expanding its applicability.

Read the paper · More papers on PaperTik