The exploration of new methods for learning in binary Boltzmann machines.

Keith Humphreys, D. Michael Titterington · 1999

Exact inference for Boltzmann machines is computationally expensive. One approach to improving tractability is to approximate the gradient algorithm. We describe a new way of doing this which is based on Bahadur's representation of the multivariate binary distribution (Bahadur, 1961). We compare the approach, for networks with no unobserved variable, to the "mean field" approximation of Peterson and Anderson (1987) and the approach of Kappen and Rodriguez (1998), which is based on the linear response theorem. We also investigate the use of the pairwise association cluster method (Tanaka and Morita, 1995). 1 INTRODUCTION The Boltzmann machine is a probabilistic network of a set of binary valued (0/1) variables fS 1 ; ::; Sn g which can be represented as an undirected graphical model. The associated probability distribution is represented by taking products of factors exp(` ij S i S j ) and factors exp(` i0 S i ): P (S 1 ; ::; Sn j`) = Z \\Gamma1 exp( n X i=1 ` i0 S i + n X i;j=...

Read the paper · More papers on PaperTik