The Discovery of Propositions in Noisy Data
Hiroshi Tsukimoto, Chie Morita · 1994
Abstract In science, the discovery of propositions in noisy data is important. This discovery is the unsupervised inductive learning with pruning. This chapter presents an efficient algorithm to induce an appropriate proposition from noisy data. A brief outline of the algorithm follows. The frequency data are transformed into a probability vector. The probability vector is transformed into a logical vector. This logical vector is then approximated by a classical logical vector. The classical logical vector is transformed into a logical proposition. This proposition is subsequently reduced to a minimum one. In this algorithm, the most original feature is the pruning method, which is executed at the approximation step. Such pruning is possible due to the transformation from a probability vector to a logical vector. The transformation is possible because (1) a proposition is represented as a vector in Euclidean space and (2) the probability vector can be transformed into a logical vector by a correspondence between the logical vector and the probability vector. Experimental results show that this algorithm performs well and also show that the propositions obtained by the algorithm basically satisfy MDL criteria. We apply this method to some problems to discover appropriate propositions or rules.