A COMPARATIVE ANALYSIS BETWEEN K-MEAN AND Y-MEANS ALGORITHMS IN FISHER'S IRIS DATA SETS.
V. Leela, K. Sakthi, R. Manikandan · 2013
Cluster analysis plays a vital role in various fields in order to group similar data from the available database. There are various clustering algorithm available in order to cluster the data but the entire algorithm are not suitable for all process .This paper mainly address with the comparative performance analysis of partition based k-mean and y-mean algorithm in Iris flower datasets. The experimental results of iris data set show that the Y-Means algorithm yields the best results in clustering and time complexity compared with k-Mean algorithm in little iteration time. Cluster analysis groups the given data objects based on only information found in the data and describes the objects and their relationships. The objective is that objects within a group be similar to one another and different from the objects in other groups. The data objects which have the maximum similarity within a group and the greater the difference between the groups are, the better or more distinct the clustering. Clustering is an effective technique for exploratory data analysis, and has found applications in a wide variety of areas. In this paper, we mainly review two algorithms k-means and y-mean algorithm. MostExisting methods of clustering can be categorized into three: partitioning, hierarchical, and grid-based and model-based methods. The k-Means and y-means are examples of partitional methods. The y-mean and k-mean are compared in the data sets of iris flower to cluster the three species of iris flower and the results are obtained in Matlab. 2. METHODOLOGY Clustering is one of the most widely performed analyses on gene expression data. Every clustering algorithm is based on the index of similarity or dissimilarity between data points. Each cluster is a collection of data objects that are similar to one another are placed within the same cluster but are dissimilar to objects in other clusters.The iris data sets are taken from three different species inorder to classify each species with common data sets. The clustering process for each algorithm differs from in order to classify the similar groups. 2.1. THE K-MEANS ALGORITHM K-Means is one of the simplest unsupervised learning algorithms used to partition the given data objects in clustering. The procedure follows a simple and easy way to classify a given data set through a certain number of clusters (assume k clusters). The main procedure is to initialize k centroids, one for each cluster groups. These centroids have to be selected carefully since their placement will always affect the end result. Finally, the k- mean algorithm aims at minimizing an objective function in the data objects, in this case a squared error function. The objective function