An Introduction to Data Mining
Daniel T. Larose, Chantal D. Larose · 2014
This chapter first provides definition to data mining. The ongoing remarkable growth in the field of data mining and knowledge discovery has been fueled by a fortunate confluence of a variety of factors such as the development of “off-the-shelf” commercial data mining software suites, and the tremendous growth in computing power and storage capacity. Automation is no substitute for human input. Humans need to be actively involved at every phase of the data mining process. The chapter then discusses cross-industry standard practice for data mining (CRISP-DM), which provides a nonproprietary and freely available standard process for fitting data mining into the general problem solving strategy of a business or research unit. It discusses four fallacies of the data mining. The chapter finally lists most common data mining tasks such as description, estimation, prediction, classification, clustering and association.