High-dimensional pattern analysis in multimedia information retrieval and bioinformatics

Yimin Wu, Aidong Zhang · Medical Entomology and Zoology · 2004

Many emerging applications, such as multimedia information retrieval and bioinformatics, require the effective analysis and retrieval of high-dimensional data. To conduct data analysis, the most extensively used method is pattern analysis (such as pattern classification and clustering). However, in high-dimensional feature spaces, pattern analysis faces the problem of the curse of dimensionality. In other words, pattern analysis can not achieve good performance in high-dimensional feature spaces, unless there exist very large volume of training data. To alleviate the curse of dimensionality, we will present an efficient feature selection method (EFS) to extract a low-dimensional feature subspace from high-dimensional feature spaces. EFS uses balanced information gain [18] to measure the contribution of each feature (for data classification); and it calculates the correlation between features with a novel extension of balanced information gain. To search the important feature subset, our approach employs a forward sequential selection algorithm to select uncorrelated features with large balanced information gain. Extensive experiments indicated that our EFS (1) can effectively alleviate the curse of dimensionality for classifying multimedia and bioinformatics data; and (2) noticeably outperforms the state-of-the-art approaches. In addition to feature selection, we will also present our extensive researches on pattern analysis in multimedia information retrieval. To improve the quality of multimedia retrieval, we presented a PatternQuest (PTQ) framework to interactively discover the useful patterns in multimedia data. PTQ is a hierarchical pattern discovery method. It resorts to our novel multiresolution pattern discovery approach (MPD) to discover useful patterns in multimedia data. MPD first organizes multimedia data repository into a hierarchical structure named data hierarchy. And then, it trains our online pattern classification methods known as adaptive random forests (ARF) and interactive random forests (IRF) to analyze multimedia data (Both ARF and IRF adapt a composite classifier known as random forests for multimedia retrieval). MPD trains the online pattern classification methods with appropriate training data from different levels of data hierarchy; and it can efficiently capture the useful patterns in multimedia data. Extensive experiments have been carried out on an image database (with 31,438 COREL images) to demonstrate the effectiveness and robustness of our method as compared against the state-of-the-art approaches.

Read the paper · More papers on PaperTik