EfficientSeg: A Self-Supervised Image-and-Text-to-Point Cloud Distillation Method Toward Annotation-Efficient Semantic Segmentation
Haoyi Zhang, Kai Wu, Weihua Li · IEEE Sensors Journal · 2025
Point cloud semantic segmentation based on LIDAR sensors is critical for intelligent vehicles to perceive and interpret their environment accurately. Traditional, fully supervised methods rely on annotations for training, which are tedious and time-consuming to obtain. Through self-supervised learning, pre-trained image features can be effectively distilled to a point cloud encoder for semantic segmentation task. However, these self-supervised methods face two challenges. Firstly, the superpixel generation suffers from many errors, leading to the “self-conflict” problem caused by the over-segmentation of semantically coherent regions. Secondly, selecting pixel features from a random view for contrastive learning is inaccurate. Objects can appear blurred and occluded within certain views. This study proposes a self-supervised method called EfficientSeg toward annotation-efficient semantic segmentation. It distills pre-trained knowledge from a vision-language model (VLM) to a point cloud encoder through the two proposed components. The multiview spatial consistent component uses a vision foundation model (VFM) to generate higher-quality superpixels and aggregate the multiview pixel features into a single representation to learn the spatial consistency. The semantic consistent component can learn the semantic consistency between the point cloud features and the text embeddings to co-optimize with the multiview spatial consistent component to reduce the impact of the “self-conflict” problem. The experiment results indicate that the EfficientSeg achieves the highest segmentation accuracy compared to other methods. The most significant accuracy improvement occurs when using only 1% annotated data for fine-tuning on nuScenes dataset. EfficientSeg achieves 47.74% mIoU, outperforming the current advanced ST-SLidR by 6.99%.