A hierarchical community of experts
Brian Sallans, Geoffrey E. Hinton, Z Ghahramani, · Library and Archives Canada (Government of Canada) · 1998
We describe a hierarchical generative model that selects from a large collection of available linear units an appropriate subset to model each observation. The selection mechanism is a corresponding network of binary units each of which gates the output of a linear unit. Inference in the binary network is intractable, but the statistics required to learn maximum-likelihood model parameters can be approximated with Gibbs sampling, even if the sampling is so brief that the Markov chain is far from equilibrium. 1 Multilayer networks of linear-Gaussian units We consider directed acyclic networks of simple stochastic units, where the units are arranged in layers. The input to a unit is the weighted sum of the activities of units in the layer above, plus a bias. In the generative model, the joint probability of all of the units in the network taking on a particular set of values, or configuration, can be factored into a product of probabilities of individual units, conditioned on the units in the layer above. The simplest unit we will consider is a linear-Gaussian unit. The probability that a linear-Gaussian unit takes on a particular value is given by a Gaussian distribution centered at the top-down prediction of the unit’s parents. The top-down prediction for unit i, denoted �yi, is the weighted sum of its parents ’ outputs, plus a bias: �yi = � j∈P a(i)