Improving Computational efficiency of Tentative Data set
Manobendu Kesari Jena · IOSR Journal of Engineering · 2012
The aim of data mining is to find novel, interesting and useful patterns from data using algorithms that will do it in a way that is more computational efficient than previous method.Current data mining practice is very much driven by practical computational concerns.In focusing almost exclusively on computational issues; it is easy to forget that statistics is in fact a core component.The Conventional decision tree classifies the data whose values are known and particular, but in this paper such classifiers handle data with tentative information.The value vagueness arises in many applications during the data collection process.Example sources of vagueness include currency conversion(Rupees to Dollar), role of packet data transfer, Daily weather report , protein design, Flight reservation, production company data due to fluctuation in raw material.With vagueness, the value of a data item is often represented not by one single value, but by multiple values forming a likelihood distribution.Rather than abstracting tentative data by statistical derivatives (such as mean and median), we discover that the accuracy of a decision tree classifier can be much improved if the "complete information" of a data item (taking into account the probability density function) is utilized .We extend conventional decision tree building algorithms to handle data tuples with tentative values and make a comparison between average based approach and distribution based approach.Since processing probability density function's is computationally more costly than processing single values (e.g., averages), decision tree construction on tentative data is more CPU demanding than that for certain data.To handle this problem, we propose a series of pruning techniques that can greatly improve edifice efficiency.