Multi-Agent Collaboration for 3D Human Pose Estimation and Its Potential in Passenger-Gathering Behavior Early Warning

Xirong Chen, Hongxia Lv, Lei Yin, Jie Fang · Electronics · 2025

Passenger-gathering behavior often triggers safety incidents such as stampedes due to overcrowding, posing significant challenges to public order maintenance and passenger safety. Traditional early warning algorithms for passenger-gathering behavior typically perform only global modeling of image appearance, neglecting the analysis of individual passenger actions in practical 3D physical space, leading to high false-alarm and missed-alarm rates. To address this issue, we decompose the modeling process into two stages: human pose estimation and gathering behavior recognition. Specifically, the pose of each individual in 3D space is first estimated from images, and then fused with global features to complete the early warning. This work focuses on the former stage and aims to develop an accurate and efficient human pose estimation model capable of real-time inference on resource-constrained devices. To this end, we propose a 3D human pose estimation framework that integrates a hybrid spatio-temporal Transformer with three collaborative agents. First, a reinforcement learning-based architecture search agent is designed to adaptively select among Global Self-Attention, Window Attention, and External Attention for each block to optimize the model structure. Second, a feedback optimization agent is developed to dynamically adjust the search process, balancing exploration and convergence. Third, a quantization agent is employed that leverages quantization-aware training (QAT) to generate an INT8 deployment-ready model with minimal loss in accuracy. Experiments conducted on the Human3.6M dataset demonstrate that the proposed method achieves a mean per joint position error (MPJPE) of 42.15 mm with only 4.38 M parameters and 19.39 GFLOPs under FP32 precision, indicating substantial potential for subsequent gathering behavior recognition tasks.

Read the paper · More papers on PaperTik