Zeta: a global method for discretization of continuous variables
Kai‐Ming Ho, Peter Dale Scott · 1997
Discretization of continuous variables so they may be used in conjunction with machine learning or statistical techniques that require nominal data is an important problem to be solved in developing generally applicable methods for data mining. This paper introduces a new technique for discretization of such variables based on zeta, a measure of strength of association between nominal variables developed for this purpose. Following a review of existing techniques for discretization we define zeta, a measure based on minimisation of the error rate when each value of an independent variable must predict a different value of a dependent variable. We then describe both how a continuous variable may be dichotomised by searching for a maximum value of zeta, and how a heuristic extension of this method can partition a continuous variable into more than two categories. A series of experimental evaluations of zeta-discretization, including comparisons with other published methods, show that zeta-discretization runs considerably faster than other techniques without any loss of accuracy. We conclude that zeta-discretization offers considerable advantages over alternative procedures and discuss some of the ways in which it could be enhanced. 1