Mining incomplete numerical data sets using C4.5 preceded by Multiple Scanning
Cheng Tian Gao, Jerzy W. Grzymala‐Busse · 2018
Noisy and incomplete data sets occur quite often in data mining tasks. If a data set is incomplete, the results extracted from the data during the data discovery phase might be inconsistent and meaningless. There are different methods for processing incomplete data sets in data mining field. Most data mining techniques can only be used on data sets composed entirely of categorical variables. To deal with continuous data, discretization is a prerequisite for most machine learning methods. Few work has been done for discretization on incomplete data. This research mainly focuses on the question whether conducting Multiple Scanning discretization as preprocessing gives better results than using C4.5 alone. Multiple Scanning is a successful discretization technique using information entropy. Previous research has shown that Multiple Scanning performs better than many commonly-used discretization algorithms. In this paper, we compare the results, in terms of error rate and tree sizes, between Multiple Scanning combined with C4.5 and C4.5 on numerical datasets with different interpretations of missing attribute values. Results shown C4.5 utilizing Multiple Scanning as preprocessing performs better than C4.5 on datasets with two types of missing data: datasets with lost values or attribute-concept values.