FEPTNet: Feature‐Enhanced Point Transformer With Channel Affinity and Graph Convolution for 3D Point Cloud Segmentation

Miaomiao Li, Kun Qi, Duo Jia, Zhiyong He · IET Computer Vision · 2026

ABSTRACT Transformer‐based networks have shown great promise in 3D point cloud semantic segmentation, yet challenges remain in effectively modelling channel dependencies and complex geometric structures. In this paper, we propose FEPTNet, a Feature‐Enhanced Point Transformer Network that addresses these limitations through two novel modules. The Channel Affinity Enhancement Module explicitly captures semantic relationships between feature channels, improving inter‐channel collaboration. The Positional Encoding‐aware Differentiated Graph Convolution replaces traditional MLP‐based position encodings with a structure‐aware graph‐based approach, enabling finer spatial perception and robustness to density variations. Built on Point Transformer V2, FEPTNet enhances both semantic representation and spatial adaptability through a U‐Net‐style encoder–decoder design. Validated on the Toronto3D benchmark, it achieves superior performance with an overall accuracy of 96.13% and a mean IoU of 80.89%. Notably, FEPTNet achieves IoU gains of 12.67% for road markings, 1.87% for road, and 4.37% for natural compared with the baseline. Compared with a wide range of representative segmentation models—including PointNet++, DGCNN, RandLA‐Net, OctNet, and Point Transformer V2—FEPTNet achieves the best overall accuracy and segmentation quality on the Toronto3D urban roadway benchmark, particularly in scenes with occlusions, structural complexity, and non‐uniform point density. As all evaluations are carried out on a single dataset, we will further explore its cross‐dataset generalisation ability in future work.

Read the paper · More papers on PaperTik