PedGT: Enhancing Pedestrian Intention Prediction Using a Skeleton-Based Graph-Transformer

Muhammad Naveed Riaz, Maciej Wielgosz, Chen Ping Xie, Antonio Manuel López · 2025

Accurately predicting pedestrian crossings in front of ego-vehicles is essential for intelligent transportation systems (ITS) to enhance road safety. Many existing approaches rely on multiple input modalities, such as scene images, segmentation maps, and trajectory data, which introduce complexity and inefficiencies, thus limiting real-time applicability. To address these challenges, we propose PedGT, a graph-based transformer model that integrates a graph convolutional network (GCN) for spatial feature extraction and a transformer encoder for temporal modeling. Unlike multi-modal methods, PedGT simplifies the pipeline by utilizing only pedestrian pose keypoints and bounding box center points, achieving superior performance on two benchmark datasets. On PIE, it achieves an F1 score of 91% and a recall of 93%, surpassing the previous best of 89%, and 88% by PCPNet. On JAAD, PedGT improves F1 and recall to 70%, outperforming PedFormer's 54% and 48%. Ablation studies highlight the impact of data normalization on accuracy, while frame importance analysis identifies keyframes influencing predictions. This work demonstrates that selecting optimal inputs and leveraging an efficient spatial-temporal model enable PedGT to outperform multi-modal solutions, providing a more streamlined and effective approach for pedestrian intention prediction.

Read the paper · More papers on PaperTik