Novel approach to data discretization
Grzegorz Borowik, Karol J. Kowalski, Cezary Jankowski · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2015
Discretization is an important preprocessing step in data mining. The data discretization method involves determining the ranges of values for numeric attributes, which ultimately represent discrete intervals for new attributes. The ranges for the proposed set of cuts are analyzed, in order to obtain a minimal set of ranges while retaining the possibility of classification. For this purpose, a special discernibility function can be constructed as a conjunction of alternative cuts set for each pair of different objects of different decisions- cuts discern these objects. However, the data mining methods based on discernibility matrix are insufficient for large databases. The purpose of this paper is the idea of implementation of a new data discretization algorithm that is based on statistics of attribute values and that avoids building the discernibility matrix explicitly. Evaluation of time complexity has shown that the proposed method is much more efficient than currently available solutions for large data sets.