EFMK: Extrinsic Parameters-Free Multi-View 3D Human Skeleton Estimation

Zijian Zhang, Muqing Wu, Honghao Qi, Min Zhao · IEEE Transactions on Circuits and Systems for Video Technology · 2025

Existing multi-view 3D human pose estimation methods heavily rely on precise extrinsic calibration, which significantly restricts their practical deployment in uncontrolled environments. To address this limitation, we propose an Extrinsic Parameter-free Multi-view 3D Human Skeleton Estimation (EFMK) framework with three technical contributions. First, a Local-Global Pose Embedding scheme is proposed to simultaneously capture the fine-grained joint dependencies while establishing cross-view correspondences. Second, a Spatial-View Joint Transformer architecture is developed with three dedicated components: (1) Feature Transformation Modulation generates adaptive modulation vectors for distinct tokens to model heterogeneous relationship patterns; (2) Prior Knowledge Enhancement systematically integrates human kinematic constraints and multi-view geometric priors into attention computation through structural topology encoding; (3) Spatial-View Joint Attention implements decoupled spatial-view attention computation followed by joint distribution modeling to capture hierarchical spatial-view dependencies. Third, a Bone-wise Reprojection-based Multi-view Aggregation mechanism is introduced to consolidate multiple 3D outputs into a single, higher-quality 3D pose for practical applications. Extensive experiments on three benchmarks demonstrate that our method achieves state-of-the-art performance while maintaining a compact model size. Code and results are available at https://github.com/Z-Z-J/EFMK.

Read the paper · More papers on PaperTik