A Monocular Fisheye Video-Based 2D to 3D Pose Lift Technique with Multiperson Spatial Context Integration
Iqbal Hassan, Nazmun Nahid, Sozo Inoue · 2024
In this paper, we propose a person intrinsic scaling method for our novel spatio-temporal encoding based 2D to 3D pose lifting technique from monocular fisheye video. While 2D pose estimation has advanced significantly, it lacks depth information and cannot capture out-of-plane movements. To handle this problem, 3D pose estimation is required. This paper proposes a framework for 2D to 3D pose lifting that tackles challenges in fisheye videos. Our method addresses fisheye deformation through a person's intrinsic scaling approach and leverages a transformer architecture with spatio-temporal self attention for pose lifting. We introduce a novel per-person scaling method to handle the depth ambiguity inherent in fisheye images. In a real life care facility data set, our framework outperformed Mediapipe, a state-of-the art method for pose lifting, in 4 different type of activity poses with a mean rmse, mae and mape value of 0.082, 0.047, and 0.046 respectively. This approach shows a new direction for 2D to 3D pose lifting technique.