Multimodal Spatio-Temporal Attention Networks with Multi-Head Residual Recurrent Encoding for Human Activity Tracking

Monirul Islam Mahmud, Md Shihab Reza, Hafeza Akter · Preprints.org · 2025

Human Activity Recognition (HAR) is vital for healthcare, behavioral monitoring, and smart environments, yet robust recognition under multimodal and low-light conditions remains challenging. We introduce a benchmark multimodal WBAN dataset combining thermal and infrared imaging with accelerometer data from Raspberry Pi 4.0, covering five core activities: eating, sleeping, staying, standing, and sitting. To model cross-modal spatio- temporal dependencies, we propose MotionXNet, a hybrid network integrating CNN feature extraction, residual BiLSTM encoding, positional embeddings, and multi-head temporal attention. Ablation studies confirm the importance of each component, with attention and positional encoding proving critical for sequence alignment and fine-grained discrimination. Compared with state-of-the-art HAR models (SenTAT, DeepConvLSTM, BiLSTM, DanHAR, MobileHART), MotionXNet achieves 96% accuracy, a macro-F1 of 0.96, and ROC-AUC of 0.9986, establishing a new benchmark for multimodal HAR.

Read the paper · More papers on PaperTik