CosPoint Transformer: Enhancing 3-D Semantic Segmentation With Cosine Similarity Attention and Cross-Attention

Thien Huynh‐The, Minh-Thanh Le, Won Jae Ryu, Anh-Kiet Vo · IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing · 2025

3-D semantic segmentation is a critical process in analyzing complex 3-D point cloud data, supporting applications across autonomous driving, urban planning, and environmental monitoring. This segmentation task has seen significant advances through deep learning, with convolutional neural networks and self-attention mechanisms enhancing accuracy compared to traditional methods. In this article, we introduce the CosPoint Transformer, a novel Transformer-based neural network for 3-D semantic segmentation using superpoints. Our model builds on the superpoint transformer by incorporating two major enhancements: cosine similarity attention (CSA) and a cross-attention (CA) mechanism in the decoder. The CSA mechanism replaces traditional multihead attention by using cosine similarity calculations to refine feature representation and improve the spatial alignment of superpoints, while the CA mechanism processes outputs from each encoder stage independently in the decoder, enabling more precise attention score calculations and enhancing segmentation performance. We evaluate the CosPoint Transformer on the DALES, S3DIS, and KITTI-360 datasets. On DALES, it achieves a mean IoU of 80.2% and 97.6% overall accuracy. On the indoor S3DIS benchmark, it obtains 66.9% mIoU and 76.4% accuracy, showcasing high efficiency with only 206 K trainable parameters utilized for both DALES and S3DIS. Furthermore, on KITTI-360, it achieves top performance with 63.4% mIoU and 92.8% accuracy using just 759 K parameters. CPT consistently outperforms its baseline and demonstrates a compelling balance of high accuracy and efficiency compared to state-of-the-art methods across these diverse outdoor and indoor environments. These results highlight the CosPoint Transformer’s effectiveness and robustness as a highly efficient solution adaptable to various 3-D semantic segmentation tasks.

Read the paper · More papers on PaperTik