A Primer on Cluster Analysis by James C. Bezdek [By the Book]
Lawrence Hall · IEEE Systems Man and Cybernetics Magazine · 2018
This book focuses on four widely used basic clustering methods suitable for most unlabeled data. The algorithms are k-means, fuzzy c-means (FCM), expectation maximization (EM), and a couple for relational clustering (sequential, agglomerative, hierarchical, and nonoverlapping). There is also a discussion of how to choose the right number of clusters for each approach. This includes internal and external validity measures for the clusters obtained by applying an algorithm. The book includes more than the four basic methods when variants and competitors are noted during discussions of the methods, and it dives into clustering approaches for big data. At the end of the chapters are historical notes, exercises, and additional examples of other algorithms. To whom is this book useful? If you have read this far, it is likely you. It is especially valuable for someone just starting to use clustering who wants background information on major underlying approaches. There are many nicely illustrated examples of clustering and validation of clusters. It is also of interest to people who know about clustering and use/develop algorithms. The book has a deep, rigorous mathematical treatment of the subject and covers work in an area with which only few are completely familiar. The historical notes are likely to be informational to most. Finally, if you teach a course on clustering, this would serve as a great textbook. It includes a deep discussion of algorithms and their variations with exercises, then moves to the modern problem of clustering big-data sets. It also shows how to evaluate the resulting partition (set of clusters). To whom is this book useful? If you have read this far, it is likely you. It is especially valuable for someone just starting to use clustering who wants background information on major underlying approaches. There are many nicely illustrated examples of clustering and validation of clusters. It is also of interest to people who know about clustering and use/develop algorithms. The book has a deep, rigorous mathematical treatment of the subject and covers work in an area with which only few are completely familiar. The historical notes are likely to be informational to most. Finally, if you teach a course on clustering, this would serve as a great textbook. It includes a deep discussion of algorithms and their variations with exercises, then moves to the modern problem of clustering big-data sets. It also shows how to evaluate the resulting partition (set of clusters).