SEZ-HARN: Self-explainable Zero-shot Human Activity Recognition Network

Devin Y. De Silva, Sandareka Wickramanayake, Dulani Meedeniya, Sanka Rasnayaka · Computers & Electrical Engineering · 2026

Human Activity Recognition (HAR) using Inertial Measurement Unit (IMU) data has numerous applications in healthcare and assisted living. However, its deployment in real-world scenarios is constrained by the limited availability of labeled datasets covering diverse activities and the lack of transparency in existing models. Zero-shot HAR (ZS-HAR) addresses data scarcity by enabling recognition of unseen activities, but current approaches largely operate as black-box models. This paper proposes SEZ-HARN (Self-Explainable Zero-shot Human Activity Recognition Network), a novel IMU-based ZS-HAR framework that simultaneously performs activity recognition and generates skeleton-based video explanations. SEZ-HARN leverages auxiliary video data to construct a semantic space and produces temporally coherent skeletal motion sequences to explain its predictions. We evaluate SEZ-HARN on four benchmark datasets — PAMAP2, DaLiAc, UTD-MHAD, and MHEALTH — and compare it with state-of-the-art ZS-HAR models. SEZ-HARN outperforms or matches conventional CNN/LSTM-based ZS-HAR models on DaLiAc (76.41%), UTD-MHAD (32.52%), and MHEALTH (46.67%) while remaining within 3% of the best-performing baseline on PAMAP2. In addition, the proposed model generates realistic, semantically meaningful explanations, as shown by alignment-based metrics and user study results. These findings indicate that SEZ-HARN effectively balances recognition performance and interpretability, making it suitable for human-centric applications.

Read the paper · More papers on PaperTik