Dynamic Adaptation of Activation Function to Fine Tune Video ResNet for Fight or Non-Fight Classification
Atif Faridi, Farheen Siddiqui, Durgesh Nandan, Md Tabrez Nafis, Mohd Abdul Ahad · Revue d intelligence artificielle · 2023
The task of designing and training a 3D convolutional neural network (CNN) from scratch poses significant complexity, necessitating high levels of expertise to achieve a performance that rivals the state-of-the-art.To circumvent this, fine-tuning of neural networks has emerged as a formidable approach.This study focuses on the utilization of Video ResNet, a state-of-the-art architecture known for its proficiency in capturing spatiotemporal patterns from video data.A novel approach is proposed for the fine-tuning of the 3D CNN model (Video ResNet) that involves altering activation functions over epochs while maintaining the network weights and biases consistent.This dynamic approach was assessed under various hyperparameters, yielding encouraging results.Contrary to most studies that employ down-sampling of the temporal sequence to minimize memory requirements, this study introduces a sliding window-based approach to evade down-sampling and prevent potential information loss.The proposed methodology yielded an accuracy of 87.25% in the fight/non-fight classification on the RWF-2000 dataset, marginally surpassing the performance of the state-of-the-art model.The proposed method not only facilitates the development of a real-time video incident detection model but also addresses the issue of overfitting during training through the incorporation of adaptive dynamic activation functions.This study thus contributes to the ongoing advancements in the field of neural network fine-tuning and video data classification.