Research on the Structure and Realization of Mixture-of-Experts
Qidong Yan, Yingjie Li, Ning Ma, Fucheng Wan · 2024
The development of Large Language Model (LLM) has reached a bottleneck, and in order to solve these problems it is necessary to continue to increase the complexity of the model. The parameters of LLM become larger and larger with the practical requirements of different application scenarios, which increases the difficulty of training as well as inference. One of the approaches to deal with the question is also known as Mixture-of-Experts (MoE), which is based on gated networks. The Mixture-of-Experts is a sparse gate-controlled deep learning model that mainly consists of a set of expert models and a gated model. The basic idea of the Mixture-of-Experts is to partition the input data into multiple regions based on the task type and assign one or more expert models to the data in each region. It is able to achieve excellent performance in complex tasks and flexibly adapt to different input distributions and task scenarios. Its high sparsity improves computational efficiency, and its specialized design enhances the ability to model complex data structures.