Video Summarization using 3D CNNs: A Convolutional Approach to Spatial-Temporal Feature Extraction

Pamidisetti Kushwanth, Kundrapu Jaswanth Naidu, Gantasala Tanooj Vardhan, Vadlamudi Manvitha, Pusarla Sindhu · 2025

The advancement of videos has been accelerating quite fast and needs efficient summarization mechanisms. The proposed approach in this work is an unsupervised framework with an efficient combination of spatiotemporal transformers and dual-path attention-based networks, yielding more accurate summaries using multimodal inputs, namely Visual, Audio, and Text. By employing a more robust attention mechanism like AoA, one is ensured of keyframe selection that is precise and maintains aesthetics and narrative coherence. It achieves personalization based on clustering-based keyframe extraction and an adaptive user preference that does not require annotated datasets. Our scalable summary achieves better F1-scores than existing deep learning methods and also resolves video-length dependency problems of the reference works. It also balances summation brevity, visual quality, and narrative fidelity while maintaining applicability to video sub-genres. Comprehensive experiments on SumMe and TVSum showed the model renders high-quality and scalable video summaries, providing a reasonable baseline for large-scale video decoding applications. Index Terms—Video summarization, Convolutional Neural Networks, Kaiming Initialization, spatial-temporal feature extraction, Binary Cross- Entropy Loss.

Read the paper · More papers on PaperTik