Multi-Layer Multi-Instance Learning for Video Concept Detection
Zhiwei Gu, Tao Mei, Xian‐Sheng Hua, Jinhui Tang, Xiuqing Wu · IEEE Transactions on Multimedia · 2008
This paper presents a novel learning-based method, called “multi-layer multi-instance (MLMI) learning,” for video concept detection. Most of existing methods have treated video as a flat data sequence and have not investigated theintrinsic hierarchy structureof the video content deeply. However, video is essentially a kind of media with ML structure. For example, a video can be represented by a hierarchical structure including, from large to small,shot,frame, andregion, where each pair of contiguous layers fits the typical MI setting. We call such a ML structure and the MI relations embedded in the structure as the MLMI setting. In this paper, we systematically study both ML structure and MI relations embedded in video content by formulating video concept detection as a MLMI learning problem. Specifically, we first construct a MLMI kernel to simultaneously model such ML structure and MI relations. To deal with theambiguity propagationproblem which is introduced by weak labeling and ML structure, we then propose a regularization framework which takeshyper-bagprediction error, sublayer prediction error, inter-layer inconsistency measure, and classifier complexity into consideration. We have applied the proposed MLMI learning method to concept detection task over TRECVid 2005 development corpus, and report better performance to vector-based and the state-of-the-art MI learning methods.