Mitigating Catastrophic Forgetting in Continual Learning for Natural Language Processing Tasks
J. Ranjith, Dr. Santhi Baskaran · Nanotechnology Perceptions · 2024
Catastrophic forgetting remains a critical challenge in continual learning scenarios for Natural Language Processing (NLP) tasks, to which this research paper provides solutions. When neural networks forget previously learnt tasks after learning new ones, this type of forgetting is termed catastrophic. However, this phenomenon is highly detrimental in NLP because linguistic data are diverse and complex. The paper introduces a novel multi-faceted approach to mitigate catastrophic forgetting, combining five key components: We show the combination of Linguistically-Informed Elastic Weight Consolidation (LI-EWC), Dynamic Architecture Expansion with Pruning (DAE-P), Task-Specific Attention Mechanisms (TSAM), Hierarchical Knowledge Distillation (HKD) and Semantic Memory Replay (SMR). These components cooperate so that one does not lose knowledge from previous tasks but can focus on the new tasks to the best effect. We evaluated the approach on a set of diverse NLP tasks, including text classification, named entity recognition, question answering and sentiment analysis. We show significant performance gains versus current state-of-the-art techniques in the average reduction of forgetting across all tasks (27%) and an overall improvement in task performance (19%). On the first task, the method can learn four subsequent tasks while preserving 94% of the original performance, compared with 86% using standard EWC and 72% using naive fine-tuning. Finally, although the performance improved, the model size grew by merely 15% after learning all the tasks. We ran an ablation study and found that the most critical performance components were the LI EWC and HKD components. Analysis by task is also done where the outperformance is consistent across all tasks (5% – 16% over second best) and varies in proportion to the task. The research opens the door to more adaptive, flexible AI systems that can address a broad spectrum of language understanding problems in natural settings, responding to the increasing demand for continual learning for NLP.