Collective organization in an adaptative mixture of experts

Vincent Vigneron, Christine Fuchen, Jean‐Marc Martinez · HAL (Le Centre pour la Communication Scientifique Directe) · 1996

The use of multiple neural models have attracted much interest recently as predictive models in system identification and control in the statistic and in the connectionist communities. In this paper, we shall restrict our attention to one particular form of density estimation called a {\em mixture model}. As well as providing powerful techniques for density estimation, mixture models find important applications in techniques for conditional density estimation, in the technique of soft weight sharing and in mixture-of-experts model. By example, the hierarchical mixture-of-experts of Jacobs \& Jordan have a powerful representational capacity and ability to handle with any multi-input multi-output mapping problem, yelding significantly faster training through the use of Expectation Maximisation algorithm. The general approach consists to divide the problem into a series of sub-problem and assign a set of 'experts' to each sub-problem. In this paper we design a new approach coupling EM algorithm (non-parametric) and Backpropagation rule (parametric) for a supervised/unsupervised segmentation of data originating from different unknown sources. We present this method as a feasible approach for learning any inverse mapping of causal systems, since Least-squares approach often leads to extremely poor performance if the image of an input is a non-convex region in the output. But in contrast to mixture-of-experts architecture, the competitions depend on the relative performance of their experts-networks, not on the input. This reasonning leads to a somewhat Bayesian version of the 'decision' by postulating a refinement of the data. the Bayesian context is presented as an alternative to the currently used frequency approach, which does not offer clear, compelling criterion for the design of statistical methods. This approach is 'non-destructive' in the sense that it does not force the data into a possibly inappropriate representation. As a simple illustration of this problem, we consider a problem of control of uranium enrichment by spectrometry-$\gamma$ .

Read the paper · More papers on PaperTik