Decoupling Mixture-of-experts Routing from Gradient Noise: A Framework for Structured Specialization and Soft Generalization Toward Robust and Efficient Inference

Bakary Badjie, José Cecílio, António Casimiro · Expert Systems with Applications · 2026

Mixture-of-Experts (MoE) models enhance deep learning scalability by activating only a subset of specialized subnetworks per input. However, conventional MoEs often suffer from unstable expert specialization and biased routing due to entangled gradient updates and unstructured input handling. To address these problems, this paper proposes SEAS-GMoE ( S tructured E xpert A ssignment and S upervised G ating M ixture o f E xperts), a hybrid framework that decouples expert routing from gradient noise through a double-stage feature clustering and semantic pseudo-labeling mechanism. SEAS-GMoE employs a K-means and KNN-based cluster refinement technique, followed by a Siamese Neural Network (SNN) that assigns semantic pseudo-labels to clustered features. This leads to generating interpretable and semantically coherent clusters that are explicitly mapped to dedicated experts. A supervised managing (gating) network learns these mappings through a bidirectional training process, reinforcing both expert specialization and routing reliability. In the bidirectional training process, as expert specialization improves, the gating supervision signal is simultaneously enhanced, and vice versa. During inference, soft expert routing is applied through the managing network’s probabilistic output. This enables flexible aggregation of expert(s)’ predictions on new inputs. Experimental results on GTSRB, MNIST, and CIFAR-10 show that SEAS-GMoE outperforms reimplemented V-MoE and dense baselines, achieving up to 4.1% higher accuracy, 60% lower inference latency, and 11% greater specialization stability. These results confirm SEAS-GMoE’s effectiveness for robust, interpretable, and efficient inference.

Read the paper · More papers on PaperTik