Integrating inflated 3D convnet and temporal convolutions for recognizing human actions
M. Kiruthika, M. Phani Sushanth, K U Sreya · 2025
Human action recognition (HAR) is a crucial aspect of computer vision, with applications ranging from sports analysis to video monitoring. Efficiently capturing both spatial and temporal data is essential for identifying complex activities within video sequences. To identify complex activities in video sequences, it is imperative to capture both spatial and temporal data efficiently. In the proposed method, we offer a hybrid method that combines the benefits of inflated 3D convolutional (I3D) and temporal convolutional network (TCN) architectures to achieve accurate and dependable action detection. While the I3D model directly extracts temporal and spatial properties from video frames, leveraging the success of 2D CNNs into 3D, in contrast, TCNs specialize in exploiting causal convolutions to simulate remote temporal relationships. By integrating these approaches, methodologies, our method aims to address their individual limitations, offering a more nuanced understanding of action dynamics in video data.