Voxel-Based Multi-Person Multi-View 3D Pose Estimation in Operating Room
Junjie Luo, Shuxin Xie, Tianrui Quan, Xuesong Ren, Yubin Miao · Applied Sciences · 2025
The localization and pose estimation of clinicians in the operating room is a critical component for building intelligent perception systems, playing a vital role in enhancing surgical standardization and safety. Multi-view, multi-person 3D pose estimation is a highly challenging task—especially in the operating room, where the presence of sterile clothing, occlusion from surgical instruments, and limited data availability due to privacy concerns exacerbate the difficulty. While voxel-based 3D pose estimation methods have shown promising results in general scenarios, their performance is significantly challenged in surgical environments with limited camera views and severe occlusions. To address these issues, this paper proposes a fine-grained voxel feature reconstruction method enhanced with depth information, effectively mitigating projection errors caused by reduced viewpoints. Additionally, an attention mechanism is integrated into the encoder–decoder architecture to improve the network’s capacity for global information modeling and enhance the accuracy of keypoint regression. Experiments conducted in real-world operating room scenarios, using the Multi-View Operating Room (MVOR) dataset, demonstrate that the proposed method maintains high accuracy even under limited camera views and outperforms existing state-of-the-art multi-view 3D pose estimation approaches. This work provides a novel and efficient solution for human pose estimation (HPE) in complex medical environments.