University of Central Florida at TRECVID 2008 Content Based Copy Detection and Surveillance Event Detection
Omer Bilal Orhan, Jason E. Hochreiter, Jonathan Poock, Qinfeng Chen, Ajay Chabra, Mubarak A. Shah · TRECVID · 2008
In this paper, we describe our approaches and experiments in content-based copy detection (CBCD) and surveillance event detection pilot (SEDP) tasks of TRECVID 2008. We have participated in the video-only CBCD task and four of the SEDP events. The CBCD method relies on sequences of invariant global image features and efficiently matching and ranking of those sequences. The normalized Hu-moments are proven to be invariant to many transformations, as well as certain level of noise, and thus are the basis of our system. The most crucial property of proposed CBCD system is that it relies on the sequence matching rather than independent frame correspondences. The experiments have shown that this approach is quite useful for matching videos under extensive and strong transformations which make single frame matching a challenging task. This methodology is proven to be fast and produce high F1 detection scores in the TRECVID 2008 task evaluation. We also submitted four individual surveillance event detection systems. “Person-Runs”, “Object-Put”, “OpposingFlow” and “Take-Picture” are the four selected events. The systems rely on low level vision properties such as optical flow and image intensity as well as heuristics based on a given event and context. The Computer Vision Lab team at University of Central Florida (vision@ucf) participated in the video-only content based copy detection and surveillance event detection pilot tasks. We submitted one base run for content based copy detection and one base run for the surveillance events of “Person-Runs”, “Opposing-Flow”, “TakePicture” and “Object-Put”. We used C++ and the OpenCv library to implement our systems. Based on the evaluation results, our copy detection approach performed very well. In terms of F1 scores, we achieved the highest results for nine out of the ten transformation types. Comparatively, our event detection results were not so significant, but still reasonable. For the last transformation, even though we could achieve high recall, we chose not to incorporate it to our base run due to low precision in detection. 1. Content Based Copy Detection (Video-Only) Our CBCD solution is an efficient computer vision system composed of 3 main parts: feature extraction & indexing, matching & retrieval and picture in picture detection & extraction. Our system is implemented using C++, OpenCv and a Java platform. Thus, it is significantly faster than comparable Matlab implementations. We have submitted one base run: VisionUCF.v.base: Result of the proposed computer vision system. The last transformation type is not included in this run because of its low precision. We rely on global image features densely extracted from the target videos for indexing. More specifically, we used Normalized Hu-Moment invariants (NHMI) taken from every frame. Since we are extracting minimal information from each frame, an exact frame match becomes irrelevant. While we lose this capability, the overall sequence matching becomes more robust and efficient. This is the main novelty of our approach, since the frame matching itself is complex under a variety of different transformations and trying to match a sequence of frames makes the problem more complicated. Our approach reduces the dimension of the search space by sacrificing the information of individual frames but keeping the video sequence data. Because we have near 20 million frames in our search corpus, we needed to use serialized Java objects to efficiently save and maintain the indexed data.