A More Rational Model of Categorization - eScholarship
Thomas L. Griffiths, Danielle Navarro, Adam N. Sanborn · Proceedings of the Annual Meeting of the Cognitive Science Society · 2006
A More Rational Model of Categorization Adam N. Sanborn ([email protected]) Department of Psychological and Brain Sciences, Indiana University, Bloomington, IN 47405, USA Thomas L. Griffiths (tom [email protected]) Department of Cognitive and Linguistic Sciences, Brown University, Providence, RI 02912, USA Daniel J. Navarro ([email protected]) School of Psychology, University of Adelaide, Adelaide SA 5005, Australia Abstract rich structures that emerge as we learn more about our environment. Accordingly, a crucial aspect of the model is the method by which stimuli are assigned to clusters. There are two steps involved in defining any ratio- nal model of cognition: first, identifying the underlying computational problem, and second, showing how peo- ple might solve that problem given cognitive constraints. When Anderson (1990, 1991) introduced the RMC, he assumed two strong cognitive constraints: that stimuli are assigned to clusters sequentially, and that these as- signments are fixed once they are made. He then intro- duced an algorithm for assigning stimuli to clusters that satisfied these constraints. However, without considering other algorithms for solving this problem it is impossi- ble to tell whether the model’s predictions result from casting categorization as a Bayesian inference about the clustering of objects, or from the assumptions about the way in which people perform this inference. Connections between the RMC and nonparametric Bayesian density estimation provide a way of defining alternative algorithms for assigning stimuli to clusters. The two algorithms we present here both asymptoti- cally approximate the Bayesian posterior distribution over assignments of stimuli to clusters, thus resulting in a “more rational” model of categorization. With these algorithms, the assumptions of the statistical model used in the RMC are no longer conflated with cognitive con- straints, and can be tested directly. These algorithms also suggest a novel class of psychologically plausible procedures for performing approximate Bayesian infer- ence under a range of cognitive constraints. To evalu- ate these ideas, we examine how well the different algo- rithms approximate both the true posterior distribution and human judgments using two data sets: the classic experiment of Medin and Schaffer (1978) and results on order-sensitivity reported by Anderson (1990). The rational model of categorization (RMC; Anderson, 1990) assumes that categories are learned by cluster- ing similar stimuli together using Bayesian inference. As computing the posterior distribution over all assign- ments of stimuli to clusters is intractable, an approxi- mation algorithm is used. The original algorithm used in the RMC was an incremental procedure that had no guarantees for the quality of the resulting approxima- tion. Drawing on connections between the RMC and models used in nonparametric Bayesian density esti- mation, we present two alternative approximation al- gorithms that are asymptotically correct. Using these algorithms allows the effects of the assumptions of the RMC and the particular inference algorithm to be ex- plored separately. We look at how the choice of inference algorithm changes the predictions of the model. Category learning is one of the most extensively stud- ied aspects of human cognition, with computational models that range from strict prototypes (e.g., Reed, 1972) to full exemplar models (e.g., Medin & Schaf- fer, 1978; Nosofsky, 1986). Recent work has emphasized the “rational” statistical basis of these models (Ashby & Alfonso-Reese, 1995), noting that prototype and ex- emplar models correspond to different approaches to the “density estimation” problem, in which one infers the probability distribution over stimuli associated with a category. These connections help to explain the suc- cess of the models and suggest new directions in which they can be extended. In this paper we discuss the sta- tistical foundations of Anderson’s (1990) model, one of the first explicitly rational approaches to category learn- ing, articulating the relationship between this model and nonparametric Bayesian density estimation. Recogniz- ing this relationship provides the opportunity to explore variations on the original model. The rational model of categorization (RMC; Ander- son, 1990, 1991) accounts for many of the basic cate- gorization phenomena, although it is not without flaws (e.g., Murphy & Ross, 1994). The RMC uses a flexible representation that can interpolate between prototypes and exemplars by clustering stimuli into groups, 1 adding new clusters to the representation as required. When a new stimulus is observed, it can either be assigned to one of the pre-existing clusters, or to a new cluster of its own. The representation can thus grow to accommodate the The rational model of categorization According to the RMC, categorization is a special case of feature induction, in which the learner uses the observed features of a stimulus to predict its unobserved features, using the previous stimuli to guide the prediction. Since the model treats category labels as features, these labels are the obvious features to predict, but other features can be predicted as well. It is assumed that each stimulus belongs to a single cluster, and that the features of a stimulus are generated by the cluster to which it belongs. Anderson (1990, 1991) refers to these groupings of stimuli as “categories”, but since they do not necessarily correspond to the category labels we will refer to them as “clusters”.