Survey on Image Segmentation Using Different K-Mean Algorithms
Miki K. Patel, Mitula Pandya · International Journal of Scientific Research · 2012
Image segmentation is the process of partitioning a digital image into multiple segments. The goal of seg- mentation is to simplify and/or change the representation of an image into something that is more meaningful and easier to analyze. Data clustering is a technique of data mining in which, the information which is logically similar is physically stored together. This paper presents the survey on image segmentation. For that the different k-mean clustering algorithms are de- scribed. The adaptive and pillar k-mean algorithms are found more efficient as those improve the segmentation quality in aspects of precision and execution time. I. Introduction Data mining is the process of extracting the information from a data set and transforming it into an understandable structure for further use. The actual data mining task is the automatic or semi-automatic analysis of large quantities of data to extract previously unknown interesting patterns such as groups of data records, unusual records and dependencies. Data mining is pri- marily used today by companies with a strong consumer focus - financial, retail, communication, and marketing Organizations. While large-scale information technology has been evolving separate transaction and analytical systems, data mining pro- vides link between the two. Data mining software analyzes rela- tionships and patterns in stored transaction data base on open- ended user queries. Clustering is the main part of the data mining. It is the task of assigning a set of objects into groups so that the objects in the same cluster are more similar to each other than to those in other clusters. Clustering is not an automatic task, but an itera- tive process of knowledge discovery or interactive multi-objec- tive optimization that involves trial and failure. It will often be necessary to modify preprocessing and parameters until result achieves the desired properties. Clustering algorithms can be categorized based on their cluster model, as hierarchical clus- tering, centroid-based clustering, distribution-based clustering, density-based clustering etc. Connectivity based clustering, also known as hierarchical clus- tering, is based on the core idea of the objects being more re- lated to nearby objects than to farther objects. Such that, these algorithms connect objects, to form based on their distance. A cluster can be described largely by the maximum distance needed to connect parts of the cluster. These algo- rithms do not provide single partitioning of the data set, but in- stead provide an extensive hierarchy of clusters that merge with each other at certain distances. In centroid-based clustering, clusters are represented by cen- tral vector, which may not necessarily be a member of the data set. When the number of clusters is fixed to 'k' cluster, k-means clustering gives a formal definition as an optimization problem: find the k -cluster centers and assign the objects to the nearest cluster center, such that the squared distances from the cluster is minimized. The clustering model most closely related to sta- tistics is based on distribution based clustering. Clusters can then easily define as objects belonging most likely to the same distribution. A nice property of this approach is that this closely resembles the way artificial data sets are generated by sampling random objects from the distribution. In density-based clustering, the clusters are defined as areas of higher density than the remainder of the data set. Objects in these sparse areas that are required to separate clusters are usually considered to be noise and border points. II. Basic Concepts In data mining, k-means clustering is a method of cluster which aims to partition n observations into k clusters in which each observation belongs to the cluster with the nearest mean value. k-means clustering tends to find clusters of comparable spatial extent, while the expectation-maximization mechanism allows clusters to have different shapes of image.