CAIRAD: A co-appearance based analysis for Incorrect Records and Attribute-values Detection
Md Geaur Rahman, Md Zahidul Islam, Terry R. J. Bossomaier, Junbin Gao · 2012
Data pre-processing and cleansing play a vital role in data mining for ensuring good quality of data. Data cleansing tasks include imputation of missing values, and identification and correction of incorrect/noisy data. In this paper, we present a novel approach called Co-appearance based Analysis for Incorrect Records and Attribute-values Detection (CAIRAD). For a data set having incorrect/noisy values CAIRAD separates the noisy records from the clean records. It thereby produces two data sets; a clean data set and a data set having all noisy records. It also reports noisy attribute values of each noisy record. We evaluate CAIRAD on four publicly available natural data sets by comparing its performance with the performance of two high quality existing techniques namely RDCL and EDIR. We use various patterns (of noisy values) each having different noise levels. Several evaluation criteria such as error recall (????), error precision (????), F-measure, record removal ratio (??????), and area under a receiver operating characteristics curve (AUC) are used. Our experimental results indicate that CAIRAD performs significantly better (based on t-test analysis) than RDCL and EDIR.