Unsupervised Learning—Clustering Using K‐Means
Wei-Meng Lee · 2019
One of the common algorithms used for clustering is the K-Means algorithm. K-Means clustering is a type of unsupervised learning used when the coder has unlabeled data. The goal is to find groups in data, with the number of groups represented by K. The goal of K-Means clustering is to achieve the following: K centroids representing the center of the clusters; and labels for the training data. This chapter presents a simple example to show how clustering using K-Means works. It also explains how to implement K-Means using Python, and explains Scikit-learn's implementation of K-Means. The Silhouette Coefficient is a measure of the quality of clustering that the coder has achieved. It measures cluster cohesion, which is the space between clusters. The range of values for the Silhouette Coefficient is between -1 and 1. The chapter presents an example of how to calculate the Silhouette Coefficient of a point.