Effects of Fusion Techniques on Sparse Sampled Frames on 3D CNN Video Classifiers

Mohammad Rasras, Iuliana Marin, Şerban Radu, Irina Mocanu · 2025

Human action recognition (HAR) has become a wide domain in computer vision over the last decade due to its usage in various applications, such as video surveillance, healthcare, sports analytics, and entertainment. In this paper, we investigate the impact of reducing temporal captured information while enhancing the spatial one on the performance of a 3D CNN model. For this, we refer to $\mathbf{R}(\mathbf{2}+\mathbf{1}) \mathbf{D}$ as a backbone framework for our experiments. We first created a base model named M-Original that modifies the original by adding a dropout layer before the final fully connected one. We tested the new design on UCF101 and obtained 85.14% accuracy. Based on the modified model, we developed three variants that differ in the attention mechanism they adopt, namely multiheaded attention, convolutional block attention module (CBAM), and temporal convolution network (TCN). Results indicate that all variants have increased the accuracy of M-Original by around 3%. In the final experiment’s part, we combined these variants in pairs to form 3 groups based on averaging, fixed, and learnable score fusion techniques. All these variants were designed to analyze the extent of performance they can add to a spatiotemporal data constrained model. Results of testing the fused pairs on UCF101 state that the learnable score-based fusion model that combines the multiheaded-attention variant with the one adopting TCN has the highest accuracy of 90.09%.

Read the paper · More papers on PaperTik