RangeViT++: refining the early convolutional stem help ViT see better

Chunyun Ma, Xiaojun Shi, Lu Chen, Shuai Song, Yingxin Wang, Jiaxiang Hu, Xialun Yun · Physica Scripta · 2024

Abstract Achieving efficient and accurate semantic segmentation of LiDAR point clouds is a crucial fundamental technology in autonomous driving and robotics. In the paper, We have designed the convolutional stem of patch embedding before the ViT in order to enhance RangeViT and improve its performance in point cloud semantic segmentation. And we have named this improved version RangeViT++. Firstly, a Multi Residual Channel Interaction Attention Module (MRCIAM) is introduced to replace original context module of RangeViT, utilizing a multi-branch structure to separately process the various channels of the range image in order to consider their modality and data distribution differences. Secondly, the Meta-Kernel module is introduced to mitigate information loss caused by the traditional CNN’s incomplete adaptation to point cloud range images, which fail to fully exploit the inherent 3D geometric information of point clouds. Lastly, during the training process, a boundary loss is incorporated to alleviate the boundary ambiguity of different classes/objects induced by the mutual conversion between point clouds and range images. Extensive qualitative and quantitative experiments conducted on challenging SemanticKITTI and SemanticPOSS dataset have verified effectiveness of our method. Superior performance is present over baseline RangeViT, which indicates refining the early convolutional stem could improve the performance of ViT on LiDAR point cloud semantic segmentation. The source code and trained model will be available at https://github.com/mafangniu/RangeViT2.git

Read the paper · More papers on PaperTik