Classification and Prediction

Rob Sullivan · Humana Press eBooks · 2011

Assigning our input instances to one of some number of distinct classes is one of the fundamental data-mining activities, allowing us a plethora of techniques to address the question “how similar are X and Y?” Once we have a sense of what characterizes such differences (or similarities), it is natural to then ask how we can predict what will come next. There are some specific data preparation challenges that we need to consider and once we have these in mind, we can focus on some of the common methods that are used by researchers. Linear regression is by far the most common method used for prediction, but decision trees and a very simple algorithm, 1R, provide valuable insight into our datasets without a huge amount of work being necessary. As we increase the sophistication of our models, the concept of the nearest neighbor becomes more important, and, following that, we discuss some aspects of Bayesian modeling and neural networks. The AutoClass algorithm has proven itself many times, and this is included as a practical method that is readily available in several implementation technologies, allowing it to be used somewhat out-of-the-box. Feature representation using alphabet sets is an obvious application area and we discuss that next before considering the k-means method and a discussion on some of the different distance measures we can use in our classification and prediction efforts. Finally, we hone in on the question of accuracy: how we measure it, what it means, and how to improve it. We include the widely used receiver operating characteristic (ROC) technique at this point also. Accuracy is directly coupled with our ability to accurately (sic) separate instances in different classes from each other. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Read the paper · More papers on PaperTik