A Vision Transformer Based Model for Human Action Recognition

Md. Ashik Billah Fahim, Mohammad Shamsul Arefin · 2025

The sucess of transformers and revolutionary attention mechanisms have motivated this reasearch on Transformer-based model for Human Action Recognition (HAR). Our proposed method utilises the architecture of the leading Video Vision Transformer (ViViT), enhanced with efficient preprocessing techniques. The integration of Tubelet Embedding technique, which embeds video clips into tubelets while preserving the temporal dimension, which captures more detailed temporal information. Then factorised self-attention mechanism is utilised for capturing both spatial and temporal information more accurately, which makes our model even more accurate. Our proposed method's results contest the current Vision Transformer (ViT)-based models and exhibit a significant enhancement in performance compared to prior state-of-the-art CNN-LSTM techniques. Our model attained a accuracy of 99.10% on the UCF11 Action dataset, showing 0.8% improvement compared to the existing methods for Human Action Classification.

Read the paper · More papers on PaperTik