Bayesian networks for lossless dataset compression

Scott Davies, Andrew Moore · 1999

The recent explosion in research on probabilistic data mining algorithms such as Bayesian networks has been focussed primarily on their use in diagnostics, prediction and efficient inference.In this paper, we examine the use of Bayesian networks for a different purpose: lossless compression of large datasets.We present algorithms for automatically learning Bayesian networks and new structures called "Huffman networks" that model statistical relationships in the datasets, and algorithms for using these models to then compress the datasets.These algorithms often achieve significantly better compression ratios than achieved with common dictionary-based algorithms such those used by programs like ZIP.Permission to make digital or hard topics of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the lirst page.To copy othcwise, to rtzublish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee.

Read the paper · More papers on PaperTik