Distilled mid-fusion transformer networks for multi-modal human activity recognition
Jingcheng Li, Lina Yao, Binghao Li, Claude Sammut · Knowledge-Based Systems · 2025
Human Activity Recognition is an important task in many human-computer collaborative scenarios, with various practical applications. Although uni-modal approaches have been extensively studied, they suffer from data quality issues and require modality-specific feature engineering, making them neither robust nor effective enough for real-world deployment. By utilizing various sensors, Multi-modal Human Activity Recognition can leverage complementary information to build models that generalize well. While deep learning methods have shown promising results, their potential in extracting salient multi-modal spatial-temporal features and better fusing complementary information has not been fully explored. Additionally, reducing the complexity of the multi-modal approach for edge deployment is another unresolved issue. To address these issues, a Knowledge Distillation-based Multi-modal Mid-Fusion approach, DMFT, is proposed to facilitate informative feature extraction and fusion for efficiently solving the Multi-modal Human Activity Recognition task. DMFT first encodes the multi-modal input data into a unified representation. The DMFT teacher model then applies an attentive multi-modal spatial-temporal transformer module that extracts the salient spatial-temporal features. A temporal mid-fusion module is also proposed to further fuse the temporal features. Subsequently, the Knowledge Distillation method is applied to transfer the learned representation from the teacher model to a simpler DMFT student model, which consists of a lite version of the multi-modal spatial-temporal transformer module, to produce the results. Evaluation of DMFT was conducted on two public multi-modal human activity recognition datasets alongside various state-of-the-art approaches. The experimental results demonstrate that the model achieves competitive performance in terms of effectiveness, scalability, and robustness.