Data Visualization & Clustering: Generative Topographic Mapping Similarity Assessment Allied to Graph Theory Clustering

Matheus de Souza Escobar, Hiromasa Kaneko, Kimito Funatsu · ACS symposium series · 2016

Chemical systems can be discriminated in several ways. If one considers industrial data, process monitoring of chemical processes can be achieved, with the following applications: anomaly discrimination (fault detection) and characterization (fault diagnosis & identification). This chapter presents an unsupervised methodology for data visualization and clustering combining Generative Topographic Mapping (GTM) and Graph Theory (GT). GTM and its probabilistic nature highlights system features, reducing variable dimensionality and calculating similarity between samples. GT, then, generates a network, clustering samples, normal and anomalous, according to their similarity. Such assessment can be applied, however, to other data sets, such as the ones involved in drug design and discovery, focusing on clustering of molecules with similar characteristics. Two case studies are presented: a simulation data set and Tennessee Eastman process. Principal Component Analysis (PCA), Dynamic PCA and kernel PCA indexes Q and T 2, along GTM independent monitoring methodologies are used for comparison, considering supervised and unsupervised approaches. The proposed method performed well for both scenarios, revealing the potential of GTM and network based visualization and clustering.

Read the paper · More papers on PaperTik