Cooperative Optimization-Based Dimensionality Reduction for Advanced Data Mining and Visualization

De Chen, S. Hamid, Michael Dix, J. A. Quirein, L. Jacobson, M. Hollingsworth · 2008

Abstract Dimensionality reduction in advanced data mining is often a non-linear problem, and the method used to resolve the dimensionality reduction will typically include: Definition of the objective function to preserve essential information of the original data Specification of the algorithm to optimize the sample mapping from a high-dimensional (HD) space to a low-dimensional (LD) space Determination of the transformation scheme that can be applied to unseen new data. In practice, a data-driven approach that maximizes the sample pair distance match of the original data and the transformed data is usually used. But, the sample mapping optimization and new data transformation are often determined in diverse domains. This paper provides a hybrid method to determine the appropriate data transformation from an original HD space to an output LD space, usually 2D or 3D space, with minimal information loss. This method is fundamentally a cooperative optimization algorithm combining more than one computational intelligence paradigm. Basically, given the problem data in a HD space, the proposed method first applies evolutionary computation (EC) to determine the LD output values, and then, performs particle swarm optimization (PSO) to refine the results. The data conversion scheme is implemented by a neural network ensemble using the EC-PSO-derived LD outputs as training targets. This method has better capability to tackle problems of local minima and to produce robust conversion of the new data. To make the preserved information essential to the user, multi-objective fitting for advanced data exploration is embedded into this method. In this paper, both the theoretical background of the developed method and the procedures applied to high-dimensional data visualization, feature extraction, and cluster analyses are discussed. Two case studies with simulated and field data to demonstrate potential applications in reservoir characterization and predictive modeling are also presented. The results strongly justify using a cooperative optimization approach to improve data mapping (especially in handling large data sets), and suggest that the method shows promise as being an effective standard procedure to help automate high-dimensional data processing.

Read the paper · More papers on PaperTik