A Self-Attention based Transformer Architecture for Enhancing the New Media Short Video Production
Jinxin Zhao · 2025
New media short video production is flourishing but labour-intensive and low-quality content creation. Manual editing, captioning, and highlight extraction are high in time and low in personalization. This study uses novel Transformer architecture to automate major parts of video production, such as summarization, captioning, and thumbnail selection. The self-attention mechanism in Transformers can well model long-range dependencies among video frames and audio tracks. The model is trained on a huge dataset of short videos to perform tasks such as speech-to-text and scene segmentation. The efficiency in content creation is increased, and viewer engagement is promoted. The accuracy scores are used for text generation and view duration for assessing video performance. The suggested transformer models achieve 96% accuracy on content engagement and half the editing time compared to RNN-based systems. The method allows creators and platforms to create high-quality videos at scale.