Spectral Clustering and Visualization: A novel Clustering of Fisher's Iris Data Set

David Benson-Putnins, Margaret Bonfardin, Meagan E. Magnoni, Daniel Michael Martin · SIAM Undergraduate Research Online · 2011

Clustering is the act of partitioning a set of elements into subsets, or clusters, so that elements in the same cluster are, in some sense, similar.Determining an appropriate number of clusters in a particular data set is an important issue in data mining and cluster analysis.Another important issue is visualizing the strength, or connectivity, of clusters.We begin by creating a consensus matrix using multiple runs of the clustering algorithm k-means.This consensus matrix can be interpreted as a graph, which we cluster using two spectral clustering methods: the Fiedler Method and the MinMaxCut Method.To determine if increasing the number of clusters from k to k + 1 is appropriate, we check whether an existing cluster can be split.Finally, we visualize the strength of clusters by using the consensus matrix and the clustering obtained through one of the aforementioned spectral clustering techniques.Using these methods, we then investigate Fisher's Iris data set.Our methods support the existence of four clusters, instead of the generally accepted three clusters in this data.Key words.cluster analysis, k-means, eigen decomposition, Laplacian matrix, data visualization, Fisher's Iris data set AMS subject classifications.91C20, 15A18 1. Introduction.Clustering is the act of assigning a set of elements into subsets, or clusters, so that elements in the same cluster are, in some sense, similar.For many, the internet is a tool used to do everything from shopping to paying bills.One can shop for clothes, groceries, movies, and more.A common theme throughout these websites is the product suggestions that appear when you buy or view an item.These product suggestions form one of the many applications of data mining and cluster analysis.Companies such as Netflix use the concept of cluster analysis to create product suggestions for their customers.The better the suggestions, the more likely the customer is to buy products.Cluster analysis can be applied to many areas including biology, medicine, and market research.Each of these areas has the potential to amass large amounts of data.There are dozens of different methods used to cluster data, each with its own shortcomings and limitations.One significant problem is determining the appropriate

Read the paper · More papers on PaperTik