Regularization and error bars for the mixture of experts network
Viswanath Ramamurti, Joydeep Ghosh · Proceedings of International Conference on Neural Networks (ICNN'97) · 2002
The mixture of experts architecture provides a modular approach to function approximation. Since different experts get attuned to different regions of the input space during the course of training, and data distribution may not be uniform, some experts may get over-trained while others are undertrained. This leads to overall poorer generalization. In this paper, we show how regularization applied to the gating network improves generalization performance during the course of training. Secondly, we address the issue of estimating the error bars for network prediction. This is useful to estimate the range of probable network outputs for a given input especially in performance critical applications. Equations are derived to estimate the variance of the network output for a given input. Simulation results are presented in support of the proposed methods which substantially improve the effectiveness of mixture of experts networks.