SPID4.7: Discretization Using Successive Pseudo Deletion at Maximum Information Gain Boundary Points
Somnath Pal, Himika Biswas · 2005
Discretization is the process of converting the continuous attributes of the database into discrete ones in order to apply some classification algorithms. This is an important problem in developing generally applicable methods in machine learning and data mining for classification and prediction. This paper intorduces a new technique for discretization based on successive pseudo deletion of instances to reduce the conflicting instances, i.e., by reduction of noise in the database. Such successive pseudo deletions in the database are performed by introducing threshold points on maximum information gain boundary points of the continuous attributes. Our empirical experiments show that the state of the art algorithms for learning, such as CN2, C4.5, Naive-Bayes and RISE give improvement in performances with the discretized output from our method than outputs with other state of the art discretization algorithms.