Semantic Segmentation from 3D LiDAR Data Using Visual Large Language Model (VLLM)

Kaan Er, Ryozo Kiyohara, Ban-Hoe Kwan, Choon‐Hian Goh, Joo-Ling Loo, Danny Wee-Kiat Ng · 2025

3D LiDAR sensors are widely utilized in robotics and autonomous systems for depth sensing and object detection. However, existing state-of-the-art point cloud detection models often require extensive fine-tuning to adapt to different environments, which may limit their scalability and practicality. To address this limitation, we proposed a novel general-purpose object recognition pipeline that integrates point cloud clustering, surface reconstruction and Visual Large Language Models (VLLMs). The process begins with acquiring point cloud data from a 3D LiDAR sensor, followed by clustering to segment individual objects. A surface reconstruction algorithm is then applied to generate structured 3D meshes, which are rendered into 2D images. These images are passed to a pretrained VLLM for object recognition using either zero-shot or prompted queries. Experimental validation on three distinct objects demonstrated successful recognition in all cases without the need for task-specific retraining. This approach offers a flexible and scalable solution for 3D object recognition, significantly reducing training costs and enhancing adaptability across diverse operational settings.

Read the paper · More papers on PaperTik