Adversarially Diversified Rehearsal Memory (ADRM): Mitigating Memory Overfitting Challenge in Continual Learning
Hikmat Khan, Ghulam Rasool, Nidhal Carla Bouaynaya · 2024
Continual learning (CL) focuses on learning non-stationary data distribution without forgetting previous knowledge. Rehearsal-based approaches are commonly used to combat catastrophic forgetting. However, these approaches suffer from a problem called “rehearsal memory overfitting”, where the model becomes too specialized on limited memory samples and loses its ability to generalize effectively. As a result, the effectiveness of the rehearsal memory progressively decays, ultimately resulting in catastrophically forgetting the learned tasks.To address the memory overfitting challenge, we introduce the Adversarially Diversified Rehearsal Memory, or ADRM, a novel method designed to enrich memory sample diversity and bolster resistance against natural and adversarial noise disruptions. ADRM employs the Fast Gradient Sign Method (FGSM) to introduce adversarially modified memory samples, achieving two primary objectives: enhancing memory diversity and fostering a robust response to continual feature drifts in memory samples.We conducted extensive experiments on the CIFAR10 dataset and found that ADRM outperforms several existing CL approaches and performs comparable to state-of-the-art methods. Additionally, we demonstrated ADRM’s ability to enhance CL model robustness under natural and adversarial conditions by using CIFAR10-C and adversarially perturbed CIFAR10 datasets.Our contributions are as follows: Firstly, ADRM addresses overfitting in rehearsal memory by employing FGSM to diversify and increase the complexity of the memory buffer. Secondly, we demonstrate that ADRM mitigates memory overfitting and significantly improves the robustness of CL models, which is crucial for safety-critical applications. Finally, our detailed analysis of features and visualization demonstrates that ADRM mitigates feature drifts in CL memory samples, significantly reducing catastrophic forgetting and resulting in a more resilient CL model. Additionally, our in-depth t-SNE visualizations of feature distribution and the quantification of the feature similarity further enrich our understanding of feature representation in existing CL approaches. Our code is publically available at https://github.com/hikmatkhan/ADRM.