Surveillance Video Analytics Using LLM-Based Summarization and Captioning
H Yamuna, A. H. Shanthakumara · 2025
Surveillance video analytics is essential today because it reduces human work, enhances security and improves efficiency.This paper introducing a system that utilizes transformer-based models BART and BLIP; these are trained using datasets BookCorpus, English Wikipedia and Conceptual Captions, COCO Captions respectively.The System generates text-based summaries and captions for key frames extracted from surveillance videos, which is particularly useful for identifying unusual activity in surveillance applications.The system extracts key frames from input surveillance video, generates captions for each frame, and summarizes the content.These summaries provide insights into key activities, security incidents, and key observations, making it easier for security personnel to review extensive video footage efficiently.BART and BLIP models outperform traditional video analysis by leveraging deep learning and these models are scalable, adaptable thus making the system more efficient for the video analysis applications.The results include a summary of activities in the video, performance metrics such as ROUGE scores for summary quality and visualizations of the processing times across different stages of the analysis pipeline.