The statistical mechanics of Bayesian model selection

Glenn Marion · ERA · 1996

and outlook viii performed using tools from statistical mechanics.In the following pages we introduce the type of neural network models we will be concerned with, describing how they can learn from examples and introducing some terminology along the way.We briefly review the recent developments in the field of artificial neural networks mentioning some notable successes in the use of these networks 1 experience.Despite great efforts, humanity has not yet been able to mimic these natural learning systems using conventional computational approaches and hence the idea behind artificial neural networks is to borrow from nature's elegant solutions.Thus, we construct models based on a network of simple neuron-like elements connected by synaptic-like weights and adapt these connections in the light of experience, that is mould the model to the data.Indeed, in the attempt to recreate some rudimentary features of natural learning systems, this approach has proved useful and, as we shall see, its novelty has lead to many developments in recent years.However, it should be stressed that the models examined in this thesis bear only passing resemblance to real neural systems. Learning from examplesThe problem of learning from examples is and has been studied in many disciplines.For example, in statistics where it is variously known as statistical inference, or nonparametric inference (in the case of neural network type models), regression, interpolation and classification (see, for example, Ripely (92) on the relation between neural networks and statistics).In the field of artificial intelligence, one refers to machine learning (e.g.Valiant (84)), whilst the problem has also been studied in the areas of speech recognition and image restoration (e.g.Geman and Geman (84)).Indeed, in comparison to these techniques, neural networks are relative new corners and should be considered as one of many possible approaches to the problem at hand.In general, the concept of learning from examples can be applied to many diverse problems but, in this thesis we focus on the problem of learning a rule which is an input-output mapping.In this problem we are supplied with a set of data by a friendly experimentalist who has gathered the data in one, or a series of experiments.In general, we note that the data will not be wholly reliable and we must take this into account in the modelling process.The observations are in the form of data pairs consisting of a set of attributes (inputs)and a set of properties (outputs) of the particular system under scrutiny.Our task is to construct an empirical model using this data which allows us to predict future values of the system properties from the attributes.That is, we wish to infer the rule relating the two; the mapping from the attributes to the properties, the inputs to the outputs.If no such rule exists, then we are wasting our time, but experience and indeed the history of science, shows that very often it does.If we are successful, in future when we measure the system attributes we will be able to predict the associated properties.This assumes that the relationship between attributes and properties does not change with time, or at least varies slowly, an assumption we will make throughout this thesis.The more general case of prediction in, so-called, non-autonomous systems is however, a growing field tackled in the artificial neural networks community by recurrent nets (see e.g.Williams and Zipser (89)).Nonetheless, the case studied here does not preclude application to dynamical systems where the attributes could represent the system at some time, t, and the properties represent the system after some fixed interval, öt; only the mapping between states at time t and t + öt must remain constant.In this latter case the time interval, 8t, will clearly affect the mapping of the attributes to the properties but could also be included as an attribute.To summarise then, in this thesis we will be concerned with the problem of learning an input-output mapping which History Interest in artificial neural networks as an alternative computational paradigm dates back to the mid-forties (McCulloch and Pitts (43)).However, the current resurgence resulted from two developments in the early to mid 1980s.A thorough account of this history is to be found in the excellent introductory text by Hertz et al.(91) and we 'spoken' by a speech synthesiser.In fact, this network learned the task to a reasonable degree but does not perform as well as the commercially available package, DECtalk, which is based on linguistic rules painstakingly elucidated over many years.However, in comparison, NETtalk achieves remarkable performance for the relatively small effort it required.We can perhaps begin to see why neural networks have been termed the second best way to do anything; the best way involving a detailed and perhaps elusive units, to be used.Linked to these issues is the crucial question of generalization.That is, given a data base of examples we want to train our student such that it will be able to generalize to situations not included in these examples, since it is of little benefit to simply reproduce the training examples themselves.Indeed, as Wolpert (92) has pointed out a simple look up table would suffice.We formally define measures of generalization performance later.In fact, in the examples discussed above the training algorithms and architectures used do critically affect the generalization ability.An intuitive understanding of why this is so can be gained from considering the common It is clear that whilst the posterior may be useful in determining the model parameters it can not be directly used to compare different models.This is because the parameters of two models with different architectures have no common interpretation.Thus, Box (80) advocates the use of the posterior to determine the model parameters whilst he argues that models themselves must be compared on the basis of the of such a measure iswhich is the square difference between the prediction, y, of some model and a sample of the output of the true teacher, drawn from P(yt I x).If the model prediction y was simply a sample from the predictive distribution P(y3 I x, V, M 2 ), namely y5 (x), then the error measure (1.10) would be equivalent to that defined by Hansen (93).However, in this thesis we will take the model output y p to be defined by equation (1.9).An Gelfand and Dey (94), Gelfand et al.(94), Neal (92) and Neal (93)).The second approach, known as Probably Almost Correct, PAC for short, derives from the computer science and artificial intelligence communities ( see Engel (94) and Anthony (95) for reviews).The PAC approach allows one to bound the generalization error of a particular model in terms of the sample complexity, p, and a model dependent quantity known as the Vapnik-Chervonenkis (V-C) dimension ( see e.g.Valiant (84), Baum and Haussler (89), Vapnik and Chervonenkis (71)).In particular, results in this where generalization curves can be calculated for general MLPs (see e.g.Saad and Solla (95a)(95b)) Although, it should be noted that this has not yet been achieved for arbitrary numbers of hidden units so that the universal approximation theorems, we mentioned in section 1.1, do not apply.Unfortunately, these online algorithms do not conveniently generate a posterior distribution and thus these exciting results are not relevant to our study.The calculations presented in this thesis are performed within this statistical mechanics framework and in particular develop the work of Hertz et al.(89), Bruce and Saad (94) and Sollich (94a).Of particular relevance to model selection problems are Chapter 2 A Statistical Mechanical Analysis of a Bayesian Inference Scheme for an Unrealizable Rule Abstract Within the Bayesian framework outlined in the previous chapter we consider a system that learns from examples.In particular, using statistical mechanical methods in the thermodynamic limit, we calculate the evidence and two performance measures, namely the generalization error and the consistency measure, for a linear perceptron trained and tested on a set of examples generated by a non-linear teacher.The learning task is said to be unrealizable because the student can never model the teacher without error even for noiseless examples.In fact, our model allows us to interpolate between the known linear case and an unrealizable, non-linear, case.A comparison of the hyper-parameters which maximize the evidence with those that optimize the performance measures reveals that, when the student and teacher are fundamentally mismatched, the evidence procedure is a misleading guide to optimizing the performance measures considered.However, consideration of the degradation in performance invoked by the evidence assignments, as compared with the optimal, demonstrates that the procedure is nonetheless remarkably robust. Chapter 4Finite Size Effects in Bayesian Model Selection and Generalization

Read the paper · More papers on PaperTik