Exploring data mining implementation

Karim K. Hirji · Communications of the ACM · 2001

Knowledge is the only factor of production that is not subject to diminishing returns. This inalienable truth is especially important in the current information age, where there is an extraordinary expansion of data generated and stored in computer databases future access. Identifying potentially useful knowledge from such databases is no trivial task and is resulting in the growing interest in data mining by both practitioners and researchers. To uncover relationship in data, statistical techniques such as factor analysis have been used in the past. Though traditional statistical techniques continue to be useful and effective for problems involving small data sets and a manageable number of variables, they run into a scalability roadblock when applied to problems where millions of records and thousands of variables exist. Data mining is thus emerging as a class of analytical techniques that go beyond statistics and aim at examining large quantities of data. The problems associated with data mining are fundamentally statistical in nature.

Read the paper · More papers on PaperTik