A Heuristic Lazy Bayesian Rule Algorithm.
Zhihai Wang, Geoffrey I. Webb · Australasian Data Mining Conference · 2002
LBR has demonstrated outstanding classification accuracy. However, it has high computational overheads when large numbers of instances are classified from a single training set. We compare LBR and the tree-augmented Bayesian classifier, and present a new heuristic LBR classifier that combines elements of the two. It requires less computation than LBR, but demonstrates similar prediction accuracy. for classification when all the attributes are mutu- ally independent given the class and the required probabili- ties can be accurately estimated from the training data. As- sume X is a finite set of instances, and A = fA1;A2;¢¢¢ ;Ang is a finite set of n attributes. An instance x 2 X is de- scribed by a vector , where ai is a value of attribute Ai. C is called the class attribute. Predic- tion accuracy will be maximized if the predicted class L( ) = argmaxc(P(cj ). Un- fortunately, unless occurs enough times within X, it will not be possible to directly estimate P(cj ) from the frequency with which each class c 2 C co-occurs with within X. Bayes' theorem provides an equality that might be used to help estimate P(cj ) in such a circumstance: P(cij ) = P(ci)P( jci) P( ) :