Learning with Advice: Exemplar-Free Adaptive Continual Learning with Mixture of Experts

Deepak Kandel, Dimah Dera · 2025

The exemplar-free continual learning (CL) trains the network without access to samples from previous tasks, making it a practical yet challenging learning approach. Recent advancements in CL include training robust feature extractors to generate task-specific latent feature representations within a multi-expert framework. While existing multi-expert approaches often employ a sparse subset of experts per task to manage computational costs, they still face limitations regarding stability and plasticity. Furthermore, due to the suboptimal design of the experts’ selection (known as the gating mechanism), there still exists room for improvement in mitigating forgetting while learning from sequences of tasks. To address these challenges, we introduce a novel mixture of expert (MoE) continual learning framework with a new gating mechanism and an adaptive knowledge-distillation-based loss function. The proposed Learning with Advice (LwA) framework adopts the Jensen-Shannon (JS) divergence as a gating mechanism to fine-tune a unique optimal expert that inherits knowledge selectively aggregated from MoE while learning new tasks. LwA leverages task-wise feature statistics in the latent space, modeled by the Gaussian mixture model (GMM), and the respective logits to maintain stability and improve plasticity. We also introduce a novel Bayesian MoE CL framework by leveraging the variational inference to build a Bayesian (probabilistic) expert backbone. Additionally, we apply the statistical shrinkage and Tukey’s transformations to the covariance matrix of latent features to ensure nonsingular statistics and enhance stability. The exhaustive empirical evaluations using challenging benchmark datasets demonstrate that LwA achieves state-of-the-art performance using both the deterministic and Bayesian backbone experts (e.g., the task-incremental accuracy of LwA is 20% higher in CIFAR100 and TinyImageNet200 with 10 Tasks), showcasing efficient knowledge inheritance from multiple experts managing stability-plasticity challenges on various continual learning scenarios.

Read the paper · More papers on PaperTik