Class-incremental learning for scene understanding
Z. Yang · 2025
Class-incremental learning (CIL) enables machine vision systems to continuously learn new tasks without forgetting previously acquired knowledge, a capability crucial for real-world applications like autonomous driving, robotics, and surveillance. The key concern in CIL lies in balancing plasticity, the ability to adapt to new tasks, with rigidity, the ability to retain prior knowledge. However, issues like catastrophic forgetting and background shift can undermine either plasticity or rigidity, and even both, while efficiency problems pose significant barriers to deploying CIL in practical systems. In order to address the above challenges, this thesis explores two subfields of CIL—few-shot learning and continual learning—in the context of three major scene understanding tasks: object detection, semantic segmentation, and panoptic segmentation. In few-shot object detection (FSOD), we address the efficiency challenges faced by current methods, such as high computational costs and slow adaptation. In Chapter 3, we propose an efficient pretrain-transfer framework (PTF), which performs on par with state-of-the-art methods without additional computational overhead. Building on this, we introduce knowledge inheritance (KI), a novel initializer that speeds up adaptation by reliably initializing novel class weights. We validate our approach on PASCAL VOC, COCO, and LVIS benchmarks, making this the first work to tackle efficiency in FSOD. In continual semantic segmentation (CSS), knowledge distillation (KD) is commonly used to prevent forgetting by transferring knowledge from the old model to the new one. However, we observe that these KD-based methods generally suffer from the novel-background confusion issue, where the model mistakes newly introduced foreground classes for the background class, or vice versa. To mitigate this confusion, in Chapter 4, we propose Label-Guided Knowledge Distillation (LGKD), which uses ground truth labels to guide the distillation process and prevent confusion. Our method outperforms existing state-of-the-art techniques on 2D benchmarks (Pascal VOC and ADE20K) and is further validated on our proposed first 3D point cloud CSS benchmark based on ScanNet, where LGKD demonstrates superior cross-modality generalization. For continual panoptic segmentation (CPS), our study reveals that existing CPS methods often suffer from efficiency or scalability issues. To address these limitations, in Chapter 5, we propose an efficient adaptation framework that incorporates attentive self-distillation and dual-decoder prediction fusion to efficiently preserve prior knowledge while facilitating model generalization. Specifically, we freeze the majority of model weights, enabling a shared forward pass between the teacher and student models during distillation. Attentive self-distillation then adaptively distills useful knowledge from the old classes without being distracted from non-object regions, which effectively enhances knowledge retention. Additionally, query-level fusion (QLF) is devised to seamlessly integrate the output of the dual decoders without incurring scale inconsistency. Our method achieves state-of-the-art performance on the ADE20K and COCO benchmarks. In conclusion, this thesis presents significant advancements in CIL for scene understanding across FSOD, CSS, and CPS, offering new solutions for creating adaptable and efficient AI systems capable of learning from continuous real-world data.