Parameter-level Mask-based Adapter for 3D Continual Learning with Vision-Language Model
Longyue Qian, Mingxin Zhang, Qian Zhang, Lin Zhang, Teng Li, Wei Zhang · 2024
As 3D object classification becomes increasingly essential in various applications, the need for models to continuously adapt to new scenarios is more crucial than ever. However, existing continual learning methods primarily concentrate on 2D tasks, making their direct application to 3D tasks challenging due to the inherent domain gap. To address this issue, we propose Parameter-level Mask-based Adapter (PMA) for 3D continual learning with Vision-Language Models. Our approach leverages the power of Contrastive Language-Image Pre-Training (CLIP) to extract rich multimodal features by projecting 3D point clouds into multi-angle depth maps, which are then processed by the 2D encoder aligned with rendered images. The core of our method is a parameter-level adapter with learnable masks that effectively integrates 3D point cloud features with 2D features. This mechanism flexibly selects the most relevant subsets of weights for each task, allowing subsequent tasks to utilize the learned weights without the need for updates, which effectively balances stability and plasticity during continual learning. Experimental results show that our method significantly outperforms existing state-of-the-art techniques.