Signature-based videos’ visual similarity detection and measurement
Saddam Bekhet · Lincoln Repository (University of Lincoln) · 2016
The quantity of digital videos is huge, due to technological advances in video capture,storage and compression. However, the usefulness of these enormous volumesis limited by the effectiveness of content-based video retrieval systems (CBVR) thatstill requires time-consuming annotating/tagging to feed the text-based search. Visualsimilarity is the core of these CBVR systems where videos are matched based on theirrespective visual features and their evolvement across video frames. Also, it acts as anessential foundational layer to infer semantic similarity at advanced stage, in collaborationwith metadata. Furthermore, handling such amounts of video data, especiallythe compressed-domain, forces certain challenges for CBVR systems: speed, scalabilityand genericness. The situation is even more challenging with availability of nonpixelatedfeatures, due to compression, e.g. DC/AC coefficients and motion vectors,that requires sophisticated processing. Thus, a careful features’ selection is importantto realize the visual similarity based matching within boundaries of the aforementionedchallenges. Matching speed is crucial, because most of the current research is biasedtowards the accuracy and leaves the speed lagging behind, which in many cases affectthe practical uses. Scalability is the key for benefiting from these enormous availablevideos amounts. Genericness is an essential aspect to develop systems that is applicableto, both, compressed and uncompressed videos.This thesis presents a signature-based framework for efficient visual similaritybased video matching. The proposed framework represents a vital component forsearch and retrieval systems, where it could be used in three possible different ways:(1)Directly for CBVR systems where a user submits a query video and the system retrievesa ranked list of visually similar ones. (2)For text-based video retrieval systems,e.g. YouTube, when a user submits a textual description and the system retrieves aranked list of relevant videos. The retrieval in this case works by finding videos thatwere manually assigned similar textual description (annotations). For this scenario,the framework could be used to enhance the annotation process. This is achievableby suggesting an annotations-set for the newly uploading videos. These annotationsare derived from other visually similar videos that can be retrieved by the proposedframework. In this way, the framework could make annotations more relevant to videocontents (compared to the manual way) which improves the overall CBVR systems’performance as well. (3)The top-N matched list obtained by the framework, could beused as an input to higher layers, e.g. semantic analysis, where it is easier to performcomplex processing on this limited set of videos.iThe proposed framework contributes and addresses the aforementioned problems,i.e. speed, scalability and genericness, by encoding a given video shot into a singlecompact fixed-length signature. This signature is able to robustly encode the shotcontents for later speedy matching and retrieval tasks. This is in contrast with thecurrent research trend of using an exhaustive complex features/descriptors, e.g. densetrajectories. Moreover, towards a higher matching speed, the framework operates overa sequence of tiny images (DC-images) rather than full size frames. This limits theneed to fully decompress compressed-videos, as the DC-images are exacted directlyfrom the compressed stream. The DC-image is highly useful for complex processing,due to its small size compared to the full size frame. In addition, it could be generatedfrom uncompressed videos as well, while the proposed framework is still applicablein the same manner (genericness aspect). Furthermore, for a robust capturing of thevisual similarity, scene and motion information are extracted independently, to betteraddress their different characteristics. Scene information is captured using a statisticalrepresentation of scene key colours’ profiles, while motion information is capturedusing a graph-based structure. Then, both information from scene and motion arefused together to generate an overall video signature. The signature’s compact fixedlengthaspect contributes to the scalability aspect. This is because, compact fixedlengthsignatures are highly indexable entities, which facilitates the retrieval processover large-scale video data.The proposed framework is adaptive and provides two different fixed-length videosignatures. Both works in a speedy and accurate manner, but with different degrees ofmatching speed and retrieval accuracy. Such granularity of the signatures is useful toaccommodate for different applications’ trade-offs between speed and accuracy. Theproposed framework was extensively evaluated using black-box tests for the overallfused signatures and white-box tests for its individual components. The evaluationwas done on multiple challenging large-size datasets against a diverse set of state-ofartbaselines. The results supported by the quantitative evaluation demonstrated thepromisingness of the proposed framework to support real-time applications.