An Overview of Efficient Data Mining Techniques
Sandeep Dhawan · 2014
Data mining is the process of discovering associations within huge data set, finding data patterns, anomalies, changes and significant statistical structures in the data. Conventional data analysis techniques involve formulating a hypothesis and then validating it against the dataset. On the other hand data mining techniques automatically detect significant patters in the data and these patterns can be used to formulate algorithms. An important consideration in mining huge data sets is that the result or the pattern identified should be valid, understandable, useful and novel [1]. Not to go without saying that data warehousing and maintaining large databases also principally rely on the efficiency of robust, intelligent and at times novel data mining techniques. Today data mining (techniques) are employed in nearly every sector of corporate industry. From music industry to films maintenance, medicine to sports there’s hardly any field of life without an input and integration of these data mining techniques. This paper focuses on presenting an overview of some of the most commonly used data mining techniques along with their applications. Techniques presented in this paper include sequence mining, clustering, classification, K nearest neighbors and association rule mining. Additionally, there’s a sample example in each case to help understand the basic working of each technique. Underlying branches, algorithms and process for each of these techniques are also given. Pseudo code for algorithms is also mentioned where required to ensure readers understanding with respective graphs. Paper also gives a brief overview some of the pre and post-processing data mining techniques.