Exploratory Data Analysis with Unsupervised Machine Learning
Altuna Akalın · 2020
This chapter focuses on using some of the machine learning techniques to explore genomics data. Clustering is a ubiquitous procedure in bioinformatics as well as any field that deals with high-dimensional data. It is very likely that every genomics paper containing multiple samples has some sort of clustering. Due to this ubiquity and general usefulness, it is an essential technique to learn. The method argument defines the criteria that directs how the sub-clusters are merged. During clustering, starting with single-member clusters, the clusters are merged based on the distance between them. The algorithm is initialized with randomly chosen k centers or centroids. In a sense, a centroid is a data point with multiple values.