Belief Propagation and Locally Bayesian Learning - eScholarship
Adam N. Sanborn, Ricardo Ferreira Louro Silva · Proceedings of the Annual Meeting of the Cognitive Science Society · 2009
Belief Propagation and Locally Bayesian Learning Adam N. Sanborn ([email protected]) Gatsby Computational Neuroscience Unit, University College London London, UK Ricardo Silva ([email protected]) Department of Statistical Science, University College London London, UK Abstract proximation embodies the compelling idea that local calcula- tions are optimal, but the global prediction is only approxi- mately optimal because information is lost in the communi- cation between regions. The approximation introduced in LBL was successful in fitting conditioning data, but the approximation itself has not been thoroughly studied. The goal of this paper is to connect the approximation used in LBL with algorithms in computer science and statistics. First we introduce LBL and compare it to the probabilistically correct updating algorithm, Globally Bayesian Learning. Next we introduce Belief Propagation and Assumed Density Filtering (ADF) and show how LBL relates to both these algorithms. Next we describe the effect of highlighting and show that ADF can produce a highlight- ing effect. Finally we show that LBL predicts a highlighting effect for alternating trials, while ADF does not. Highlighting, a conditioning effect, consists of both primacy- like and recency-like effects in human subjects. This combination of effects are notoriously difficult for Bayesian models to produce. An approximation to probabilistic inference, Locally Bayesian learning (LBL), can predict highlighting by partitioning the model into regions during learning and passing messages between these regions. While the approximation matches behavior in this task, it is unclear how LBL compares to other approximations used in Bayesian models, and what behaviors this approximations will predict in other paradigms. Our contribution is to show LBL is closely related to the statistical algorithms of Assumed Density Filtering (ADF), which simplifies calculations by assuming independence, and Belief Propagation, which identifies how to make these calculations through message passing. We propose that people use ADF to learn and show how this model can produce highlighting behavior. In addition, we demonstrate how the degrees of approximation used in LBL and ADF cause the models to make very different predictions in a proposed experimental design. Locally Bayesian Learning Keywords: machine learning; conditioning; Bayesian models; belief propagation; assumed density filtering The probabilistic model underlying LBL was a generaliza- tion of a feedforward neural network with one hidden layer. This structure and the variable names used are shown in Fig- ure 1a. The neural network was generalized from having sin- gle estimates of hidden weights, hidden nodes, and output weights by putting a distribution over the values of these hid- den variables. The generalization required different calcula- tions than a neural network to update the weights, and the correct probabilistic update was termed Globally Bayesian Learning (GBL), Bayesian approaches have often been successful in pre- dicting human data as the result of optimal behavior (e.g., K¨ording & Wolpert, 2004), but people can also produce be- havior that is very difficult to explain with Bayesian ap- proaches (Daw, Courville, & Dayan, 2008; Kruschke, 2006a, 2006b). Additionally, Bayesian approaches can be hampered by computational complexity in practical applications (An- derson, 1991). Modeling human cognition as an approximation to optimal behavior is an approach that has been used to both explain deviations from optimality as well as managing the computa- tional complexity of the solution (Gigerenzer & Todd, 1999; Kahneman, Slovic, & Tversky, 1982; Kruschke, 2006b). Many of these algorithms were invented to match human behavior, but recent work has proposed that candidate al- gorithms for human cognition could be drawn from work in computer science and statistics (Sanborn, Griffiths, & Navarro, 2006). These algorithms have been developed to efficiently produce results faithful to the full model, and can come with guarantees on the quality of the approximation. Conditioning has been a testbed for approximations to Bayesian models. Conditioning effects such as blocking and highlighting depend on the order of the stimuli. An approx- imation that has successfully fit these types of effects is Lo- cally Bayesian Learning (LBL; Kruschke, 2006b). This ap- p(W hid ,W out , y | x,t) ∝ p(t | W out , y)p(y | W hid , x)p(W out ,W hid ) (1) where the variables in the model are described in Figure 1 and p(t|W out , y) and p(y|W hid , x) were the standard linear combi- nation of weights plus a sigmoid nonlinearity. The network graph was then split into regions to imple- ment LBL, as in Figure 1b. Psychologically, LBL was meant to represent two stages: an early attentional phase that took stimulus cues from the input nodes and converted it into at- tended cues on the hidden nodes, and weights from the at- tended cues to the output. Correct probabilistic calculations were used within regions, but LBL uses messages between regions that result in an approximation to GBL. Informa- tion was passed between the two regions of LBL in two par- ticular ways. The expected value of the attentional nodes E(y|x) in the first region given the stimulus cues was passed