3DPCNet: Pose Canonicalization for Robust Viewpoint-Invariant 3D Kinematic Analysis from Monocular RGB Cameras

Tharindu Ekanayake, Constantino Álvarez Casado, Miguel Bordallo López · 2026

Monocular 3D pose estimators produce camera-centered skeletons, creating view-dependent kinematic signals that complicate comparative analysis in applications such as health and sports science. We present 3DPCNet, a compact, estimator-agnostic module that operates directly on 3D joint coordinates to rectify any input pose into a consistent, body-centered canonical frame. Its hybrid encoder fuses local skeletal features from a graph convolutional network with global context from a transformer via a gated cross-attention mechanism. From this representation, the model predicts a continuous 6D rotation that is mapped to an SO(3) matrix to align the pose. We train the model in a self-supervised manner on the MM-Fi dataset using synthetically rotated poses, guided by a composite loss ensuring both accurate rotation and pose reconstruction. On MM-Fi, 3DPCNet reduces the mean rotation error from over 20° to 3.4° and the Mean Per Joint Position Error from ∼64 mm to 47 mm compared to a geometric baseline. On TotalCapture database, fine-tuned with calibrated multi-view viewpoint variation, 3DPCNet reaches 0.3–1.3° rotation error and 40–44 mm MPJPE on held-out viewpoints and motions. Qualitative cross-modal evaluation on TotalCapture further shows that accelerations derived from canonicalized video poses correspond well to wrist IMU signals, supporting physically plausible motion analysis under viewpoint changes.

Read the paper · More papers on PaperTik