Locating and tracking an object across multiple video feeds using zero-shot object detection
Brad Killen, Audrey L. Aldridge, Paul M. Barrett, Jeremy Davis, Daniel W. Carruth, Cindy L. Bethel · 2025
Currently, searching for specific objects across multiple videos is a costly, resource-heavy task, due to a lack of either adequate computational processing power or personnel. To address these challenges, this paper presents a framework for detecting and tracking objects across multiple videos with minimal effort. By combining several computer vision tools, such as SAM (Segment Anything Model), YOLOv8 (You Only Look Once-version 8), and DINOv2 (Self-Distillation with No Labels-version 2), this framework requires minimal training across machine learning (ML) models and can ease the burden placed on users when parsing and monitoring hours of video footage. By optimizing this framework, the time, effort, and resources spent processing videos is reduced to a fraction of the time, allowing for more flexibility in the system’s application. Evaluation of the framework’s efficiency in identifying a specific object of interest (OoI) examines accuracy, speed, and how the optimization of each of these performance metrics affects resource and memory consumption.