Exploring Human Activity Recognition with Acoustic Data: A Comparative Study of CNN-LSTM, ViViT, and ResNet-Temporal Transformer Model

Alaa Humaidan, Jeny Roy, Sara Sharifzadeh, Ruchita Mehta, Andrea Tales, Joe Macinnes · 2025

This paper addresses the continuous Human Activity Recognition (HAR) problem using acoustic sensors, which finds application in aged population health and well-being monitoring. The challenging class imbalance problem has been studied using three main groups of time-series modelling strategies, including: 1) the local feature extraction based on Convolutional Neural Networks-Long Short-Term Memory Networks (CNN-LSTM), 2) feature learning based on global dependencies using Video Vision Transformer (ViViT), and 3) combination of the local and global features using ResNet-based frame-level feature extraction followed by Temporal Transformer (RNTT). The acoustic spectrograms have been augmented by adding noise, which has improved all models accuracy (83-86%). Research findings have also demonstrated the best level of resilience to noise condition using the proposed RNTT pipeline (93%), and then the ViViT model (86%).

Read the paper · More papers on PaperTik