DATA MINING: SHOULD IT BE INCLUDED IN THE 'STATISTICS' CURRICULUM?
Siva Ganesh V · 2002
Data Mining is the process of extracting knowledge from large volumes of data. In other words, data mining is the science (or art) of discovering unexpected patterns, valuable structures and interesting relationships in large and complex data, with an emphasis on large observational databases. The combination of fast computers and cheap storage makes it easier to extract useful information out of everything from supermarket buying patterns through Banking and Stock market trades to medical diagnostics and remote sensing. Data mining is widely used for achieving organisational goals and allowing investigators to go beyond simple data queries and reporting, in order to understand why things happen and to manipulate what happens in the future. What distinguishes data mining from conventional statistical data analysis is that data mining is usually done for the purpose of 'secondary analysis' aimed at finding unsuspected relationships, perhaps, unrelated to the purposes for which the data were originally collected. In other words, data mining is very much an inductive exercise, as opposed to the traditional hypothetico-deductive approach of statistics. Data mining sits at the common frontiers of fields such as Information Systems (Database management & Data warehousing), Computer Science (Artificial Intelligence, Machine Learning & Pattern Recognition), and Statistics (Data Visualisation & Modelling).