Fast and Accurate Video Analysis and Visualization of Classroom Activities Using Multiobjective Optimization of Extremely Low-Parameter Models
Venkatesh Jatla, Sravani Teeparthi, Ugesh Egala, Sylvia Celedón‐Pattichis, Marios S. Pattichis · IEEE Access · 2025
The paper considers the problem of video activity recognition in real-life collaborative classroom learning environments. Video analysis of real-life collaborative classroom learning environments faces significant challenges not encountered in current, advanced video recognition datasets. In collaborative learning environments, students are arranged in small groups where they interact within their group. Video analysis needs to deal with long-term activity recognition (of one hour or more session videos), detect multiple simultaneous activities, rapid transitions between activities, occlusions, and numerous individuals performing similar activities in the background that are not part of the group being analyzed. Developing ground truth datasets for analyzing complex video datasets is prohibitively expensive. We dramatically reduce the requirement for large ground truth datasets by creating separate, custom datasets for object detection and video activity recognition.We then introduce a separable, extremely low-parameter system for video activity recognition that can be optimally trained using the derived datasets without the need for transfer learning from larger systems trained on large datasets. We further develop an interactive WebApp for visualizing the results over long video sessions. Overall, the extremely low-parameter activity classification model uses just 18.7K parameters for each activity, requiring 136.32 MB of memory. On a moderate GPU (RTX 5000), the activity classification model runs at an impressive 4,620 (154 x 30) frames per second. Our approach uses at least 1,000 fewer parameters than several well-established methods for video recognition. Our extremely low-parameter classifiers can process 90 minutes of video in just 26 seconds. Furthermore, our models are much easier to train, they are much faster, and outperform comparable methods.