Semantic Segmentation of LiDAR Point Clouds Using Image Annotations
Johanna Nilsson · Diva portal (Dalarna University Library) · 2026
Semantic segmentation of LiDAR point clouds is an important but challenging task within the field of computer vision. Obtaining accurate ground truth labels for training is time-consuming, which has led to an interest in alternative approaches. This thesis investigates how pseudo ground truth labels for semantic segmentation in LiDAR point clouds can be generated by transferring labels from 2D images to 3D point clouds, when ground truth is absent for both images and point clouds. The focus is on unstructured outdoor environments. A method is proposed where 2D semantic image segmentation is performed using a pre-trained model on the GOOSE dataset. The image labels are transferred to LiDAR point clouds. The labels are pre-processed and the generated pseudo labels are used to train 3D models for semantic segmentation. The results using 2D image segmentation to generate pseudo labels are compared to the results using a model trained on 3D data from the GOOSE dataset. The visual results show that some incorrect labels are present in the labeled images from inference using the pre-trained model. Filtering the point clouds through 3D neighborhood majority voting improves the quality of the point cloud labels. The best-performing model reaches a test mIoU of 0.470 for four general outdoor categories, when evaluated on the test set with pseudo labels. In the predictions of the test set, all models manage to segment the trees well. However, the other categories are often mixed up. The proposed method demonstrates that transferring labels from 2D segmentation to LiDAR point clouds is a viable approach to generate pseudo ground truth labels that can be used for training, but the quality can be limited and depends highly on factors like 2D segmentation and calibration accuracy.