REVIEW OF LITERATURE ON DATA MINING

Tejaswini Abhijit Hilage, R. V. Kulkarni · 2012

Data mining is used for mining data from databases and finding out meaningful patterns from the database. Many organizations are now using these data mining techniques. In this paper authors has reviewed the literature of data mining techniques such as Association Rules, Rule Induction Technique, Apriori Algorithm, Decision tree and Neural network. This review of literature focuses on how data mining techniques are used for different application areas for finding out meaningful pattern from the database. Most of the different approaches to the problem of clustering analysis are mainly based on statistical, neural network, machine learning techniques. Bagirov et al. (4) propose the global optimization approach to clustering and demonstrate how the supervised data classification problem can be solved via clustering. The objective function in this problem is both nonsmooth and nonconvex and has a large number of local minimizers. Due to a large number of variables and the complexity of the objective function, general purpose global optimization techniques, as a rule fail to solve such problem. It is very important therefore, to develop optimization algorithm that allow the decision maker to find local minimizers of the objective function. Such deep mininizers provide a good enough description of the data set under consideration as far as clustering is concerned. Some automated rule generation methods such as classification and regression trees are available to find rules describing different subsets of the data. When the data sample size is limited, such approaches tend to find very accurate rules that apply to only a small number of patients. In Schwarz et al. (16) it was demonstrated that data mining techniques can play an important role in rule refinement even if the sample size is limited. For that at first stage methodology is used for exploring and identifying inconsistencies in the existing rules, rather than generating a completely new set of rules. K-mean algorithm lies in the improved visualization capabilities resulting from the two dimensional map of the cluster. Kohonen developed self organizing maps as a way of automatically detecting strong features in large data sets. Self organizing map finds a mapping from the high dimensional input space to low dimensional feature space, so the clusters that form become visible in this reduced dimensionability. The software used to generate the self organizing maps is Viscovery SOMine (www.eudaptics.com), which provides a colorful cluster visualization tool, & the ability to inspect the distribution of different variables across the map. The subject of cluster analysis is the unsupervised classification of data & discovery of relationship within the data set without any guidance. The basic principle of identifying this hidden relationship is that if input patterns are similar, they should be grouped together. Two inputs are regarded as similar if the distance between these two inputs is small. This study demonstrates that data mining techniques can play an important role in rule refinement, even if the sample size is limited. Leonid Churilov, Adyl Bagirov, Daniel Schwartz, Kate Smith and Michael Dally demonstrated that both self organizing maps & optimization based clustering algorithms can be used to explore existing classification rules, developed by experts and identify inconsistencies with a patient database. As the proposed optimization algorithm calculate clusters step by step and the form of the objective function allow the user to significantly reduce the number of instances in a data set. A rule based classification system is important for the clinicians to feel comfortable with the decision. Decision tree can be used to generate data driven rules but for small sample size these rules tend to describe outliers that do not necessarily generalize to larger data sets.

Read the paper · More papers on PaperTik