Hierarchical Clustering in R
Martin Vogt, Jürgen Bajorath · 2017
This chapter illustrates the usage of clustering methods in R as an example of unsupervised learning. The goal of cluster analysis of compound data sets is to generate an organization of compounds into different clusters (also called groups or communities) so that compounds within a cluster are, with respect to pre-defined characteristics or descriptors, more similar to each other than to compounds in other clusters. Two of the most popular representations of molecules for numerical analysis are numerical descriptors (with continuous value ranges) and binary fingerprint representations. This chapter explains hierarchical clustering using fingerprints and descriptors. It then explores visualization of the data sets. Fingerprints cannot be handled by standard R and custom code is required to generate distance matrices. The data sets consist of compounds active against different targets that are structurally diverse and can be easily distinguished using MACCS fingerprint-based clustering.