Real-Time Video Captioning on CPU and GPU: A Comparative Study of Classical and Transformer Models

Othmane Sebban, Ahmed Azough, Mohamed Lamrini · International Journal of Advanced Computer Science and Applications · 2025

This study proposes a scalable and hardware-adaptable approach to automatic video caption generation by comparing two architectures: a traditional encoder–decoder framework combining InceptionResNetV2 with GRU and a transformer-based model integrating TimeSformer with GPT-2. The system supports CPU and GPU deployment through a unified pipeline built on FFmpeg and ImageMagick for keyframe extraction and subtitle embedding. Experimental evaluations on the MSVD and VATEX datasets demonstrate that the TimeSformer–GPT-2 architecture significantly outperforms baseline models, particularly in GPU settings, achieving top results across BLEU, METEOR, ROUGE-L, and CIDEr metrics. This superiority is attributed to its capacity to model spatiotem-poral dependencies and generate contextually rich language. Designed for real-time operation, the system is also suitable for low-resource devices, enabling impactful applications such as assistive tools for the visually impaired and intelligent video indexing. Despite high computational demands and sequence-length limitations, the system presents promising directions for future development, including multilingual captioning, multimodal audio–visual integration, and lightweight models like TinyGPT for enhanced portability.

Read the paper · More papers on PaperTik