Large scale semantic concept detection, fusion, and selection for domain adaptive video search
Yu–Gang Jiang · 2009
This thesis investigates the problem of video search based on semantic concepts. We present approaches to handle three correlated issues that are critical to this problem: (1) how to construct an effective feature representation for semantic concept detection, (2) how to exploit semantic context to improve the detection of these concepts, and (3) how to select the most suitable concept detectors to answer user queries. In particular, as the target videos may come from different domains (genres or sources) with distinctive data characteristics, for each of the issues, we will need to cope with the domain changes. Video frames are represented by bag-of-visual-words (BoW) derived from local keypoint features, which are invariant to rotation, scale and illumination. We first conduct a comprehensive study on the representation choices of BoW, including vocabulary size, weighting scheme, stop word removal, feature selection, spatial information, and visual bi-gram. The aim is to offer practical insights in how these choices will impact the performance of BoW for semantic concept detection. We also show how to further augment the BoW representation by exploring the linguistic and ontological aspects of visual words. A visual-word ontology is constructed to hierarchically specify their hyponym relationship, which is incorporated into BoW for improved video frame representation. To exploit semantic context, we develop a novel and efficient domain adaptive semantic diffusion algorithm. Inter-concept relationship is modeled using a semantic graph, which treats concepts as nodes and the concept affinities as the weights of edges. It is then applied to refine the initial detection results through a function level graph diffusion process, aiming to recover the consistency and smoothness of the detection results over the graph. To handle the domain change