An Algorithm for Fine-Grained Content Extraction and Understanding in Short Videos
Yanqi Wan, Shuya Zhang, Yaqi Xu, K Y Zhang, Heyi Wang, Mingzheng Liu · Data · 2026
This study integrates communication theory with advanced computer vision techniques to propose a novel approach for fine-grained content extraction in short videos. Unlike methods focused on summarization or subtitle generation for longer videos, our approach emphasizes extracting detailed content and understanding the intricate narrative structure of short videos. By employing scene segmentation, similarity-based filtering algorithms, and support vector machines, the method identifies keyframes that capture precise visual details. Further, it generates semantically accurate textual descriptions using the mPLUG model, enabling an in-depth understanding of video content. Using a dataset of short videos from the cultural and tourism domain, we validated the proposed method. Experimental results demonstrate that our approach achieves high precision in identifying and understanding detailed visual elements, effectively bridging the gap between visual representation and semantic meaning. Additionally, the study explores the influence of different video content types, interference factors, and image description models on fine-grained content extraction, highlighting its potential for improving intelligent analysis of short video data.