An RST based efficient preprocessing technique for handling inconsistent data

Surekha Samsani · 2016

Data Preprocessing is an essential and primary step in the process of knowledge discovery; because the data obtained from the logs may be incomplete, noisy or inconsistent. The quality of the training data plays a vital role in the success of the data mining algorithms thus; Data Preprocessing should not be an exception in the process of knowledge discovery. The most promising attributes of the quality data includes completeness, consistency and timeliness. Mainly the existence of Inconsistent data misguides the mining algorithm and indirectly affects the performance of the data mining algorithm. This paper gives, Rough Set Theory based approach for identifying and dealing with such inconsistencies in the given dataset and its performance is tested by submitting the preprocessed data to the tree based classifier. Finally, the experimental results revealed the importance of data preprocessing.

Read the paper · More papers on PaperTik