Some Methods for Training Mixtures of Experts

Perry D. Moerland · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 1997

Introduction Recently, a modular architecture of neural networks known as a mixture of experts (ME) [8][9] has attracted quite some attention. MEs are mixture models which attempt to solve problems using a divide-and-conquer strategy; that is, they learn to decompose complex problems in simpler subproblems. In particular, the gating network of a ME learns to partition the input space (in a soft way, so overlaps are possible) and attributes expert networks to these different regions. The divide-and-conquer approach has shown particularly useful in attributing experts to different regimes in piece-wise stationary time series [20], modeling discontinuities in the input-output mapping, and classification problems [6][13][19]. The ME error function is based on the interpretation of MEs as a mixture model [12] with conditional densities as mixture components (for the experts) and gating network outputs as mixing coefficients. The purpose

Read the paper · More papers on PaperTik