Poster: Multimodal ConvTransformer for Human Activity Recognition
Syed Tousiful Haque, Anne H. H. Ngu · 2024
Recognition of human activity is crucial in emer-gency healthcare applications. Multi-modal deep learning learning algorithms are gaining attention for Human Activity Recog-nition (HAR) due to their success in various domains. The application of multi-modal learning to HAR continues to present challenges, particularly in addressing noisy data and achieving effective fusion of information from disparate modalities. We propose a new Multimodal-ConvTransformer (CT-HAR) that aims to efficiently extract and fuse complementary spatial and temporal information. Experiments conducted on the public UTD-Mhad and Berkley-Mhad datasets demonstrate significant performance enhancements, with CT-HAR achieving accuracy rates of 89.81 % and 85.69% on these datasets, respectively.