An Experimental Comparison of Methods for Handling Incomplete Data in Learning Parameters of Bayesian Networks

Agnieszka Oniśko, Marek J. Drużdżel, Hanna Wasyluk · 2002

Missing values of attributes in data sets, also referred to as incomplete data, pose difficulties in learning tasks, such as classification, data mining, or learning Bayesian network structure and its numerical parameters. Because of the predominance of incomplete data in practice, many methods have been proposed to deal with them while there are few studies that compare their performance. The Hepar II project presents an excellent opportunity to test experimentally how these methods perform on a real data set. We briefly review several popular methods for handling incomplete data and then compare them on the task of learning conditional probability distributions of a Bayesian network model, where the comparison criterion is the resulting diagnostic accuracy. While substitution of “normal” values of missing attributes seemed to perform best, we observed only a small difference in performance among the studied methods. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Read the paper · More papers on PaperTik