Advanced tools for video and multimedia mining
Christos Faloutsos, Howard D. Wactlar, Jia-Yu Pan · 2006
How do we automatically find and mine data in large multimedia databases, to make these databases useful and accessible? We focus on two problems: (1) mining patterns that summarize characteristics of a data modality, and (2) mining among multiple modalities. Uni-modal such as videos have static scenes and speech-like sounds, and cross-modal correlations like the blue region at upper part of a natural scene image is likely to be 'sky', could provide insights on multimedia content and have many applications. For uni-modal pattern discovery, we propose method AutoSplit . AutoSplit provides a framework for mining meaningful independent components in multimedia data, and can find in a wide variety of data modalities (e.g., video, audio, text, and time sequences). For example, in video clips, AutoSplit finds characteristic visual/auditory patterns, and can classify news and commercial clips with 81% accuracy. In time sequences like stock prices, AutoSplit finds hidden variables like general growth trend and Internet bubble, and can detect outliers (e.g., lackluster stocks). Based on AutoSplit, we design a system, ViVo, for mining biomedical images. ViVo automatically constructs a visual vocabulary which is biologically meaningful and can classify 9 biological conditions with 84% accuracy. Moreover, ViVo supports data mining tasks such as highlighting biologically interesting image regions, for biomedical research. For cross-modal correlation discovery, we propose MAGIC, a graph-based framework for multimedia correlation mining. When applied to news video databases, MAGIC can identify relevant video shots and transcript words for event summarization. On task of automatic image captioning, MAGIC achieves a relative improvement of 58% in captioning accuracy as compared to recent machine learning techniques.