Temporal-Spatial-Relation Former for Multi-Person Motion Prediction
Yun Zhang, Chiyu Cai, Xiaoling Luo, Ping Li, Yalan Ye · IEEE Transactions on Consumer Electronics · 2025
Multi-person motion detection remains a challenging problem due to the highly complex spatiotemporal dynamics it involves. Effective motion forecasting requires capturing both the internal joint movement patterns of individuals and the interactions among multiple individuals. While existing Transformer-based models have demonstrated remarkable performance, they predominantly focus on spatiotemporal features, often neglecting crucial relational interactions between joints. To address this limitation, we propose the Temporal-Spatial-Relation Former (TSRFormer), a novel model for multi-person 3D pose forecasting that integrates both spatiotemporal and relational features. The proposed TSRFormer consists of two complementary branches: the first branch leverages the Spatial Transformer and Temporal Transformer to extract spatial and temporal dependencies of motion joints, while the second branch employs a convolutional network to capture inter-person relationships. These branches interact and exchange information within the proposed computational mode, enabling a holistic understanding of motion dynamics from multiple perspectives. The experiments on benchmark datasets demonstrate that TSRFormer achieves competitive performance, particularly excelling in long-term motion prediction, thereby validating its effectiveness in capturing complex temporal dependencies and relational interactions in multi-person scenarios.