Unsupervised Learning and Optimization
John Tuhao Chen, Lincy Y. Chen, Clement M. Lee · 2024
The previous chapters discuss data analytic issues on input features (such as predictors) relating to output features (such as the response variable), where each observation has a response. For instance, in the analysis of clinical factors related to systolic blood pressure, the response variable is the reading of patients&s; systolic blood pressure; in the classification of up or down market trend in the coming time period, the response variable is either “bull market” or “bear market”. The model learned from the training data has a response variable intended to “supervise” the learning process by using a MSE criterion or the total probability of correct classification. However, in some data analytic problems, the response to “supervise” the learning process might not even exist. For example, in business analysis, clustering of consumer preferences helps structure the design of marketing strategies of advertising campaigns. In clinical trials and epidemiology, grouping patient symptoms helps diagnosis and prevention in public interventions. In geology, grouping on element characteristics of rock samples helps identify the main characteristic of the environment it was found. The common theme among the above mentioned applications is the lack of a response variable, due to the absence of knowledge in the experiment stage. In this chapter, we will focus on two main methods: K-means clustering and the method of principal component analysis. To briefly summarize, K-means clustering and principal component analysis are two optimization approaches in grouping a set of data.