Vision–language guided semantic-geometric transformer for memory-efficient 3D scene understanding

Licheng Liu, Yu Li, Fuyong Liu · Frontiers in Artificial Intelligence · 2026

Recent 3D Transformers have become a dominant framework for point-cloud segmentation by modeling spatial context in sparse 3D scenes. However, geometry and color alone provide limited high-level semantic cues, especially for cluttered boundary regions, visually similar objects, and long-tail categories. To address this issue, we propose a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors. Specifically, dense CLIP features extracted from multi-view RGB images are projected onto 3D points through visibility-aware alignment and view pooling, and are fused with relative geometric offsets and color cues to form semantically aware sparse voxel tokens. To better exploit the aligned CLIP semantics during local token interactions, we build on contextual relative signal encoding (cRSE) and introduce a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention. We further adapt block-wise online softmax computation to generate and consume these biases on the fly. Experiments on ScanNet, ScanNet200, and S3DIS demonstrate competitive segmentation performance, improved instance-level discrimination, and a 25.7% reduction in peak online 3D-stage training memory compared with the materialized attention implementation when cached CLIP features are used.

Read the paper · More papers on PaperTik