DeepMoE: MoE for deep non-hierarchical representation mechanisms
Hengye Wang, Jianing Wang · 2025
Sparse Mixed Expert Models (MoE) have received widespread attention for their ability to increase the number of model parameters while reducing computational overhead. However, existing sparse hybrid expert models arrange all experts in the same layer with shallow network depth, and each expert receives data with the same dimension, which can limit the expressive power of the model and make it difficult to fully capture the diversity and hierarchy of complex data. In this paper, we propose DeepMoE, an efficient multilevel hybrid expert architecture, which realizes a deep sparse hybrid expert model by fusing the advantages of residual structures and fully connected networks, analogous to the information transfer mechanism in the cerebral cortex of the human brain. Our experiments show that DeepMoE improves performance on several tasks such as VQA and ScienceQA, demonstrating its potential and application value in multimodal data processing.