KPDepth-VO: Self-Supervised Learning of Scale-Consistent Visual Odometry and Depth With Keypoint Features From Monocular Video
Changhao Wang, Guanwen Zhang, Zhengyun Cheng, Wei Zhou · IEEE Transactions on Circuits and Systems for Video Technology · 2025
Monocular visual odometry (VO) is crucial for the application of various autonomous systems. However, the inherent scale ambiguity issue in monocular methods greatly limits their performance in pose estimation. In this paper, we propose a hybrid monocular VO system named KPDepth-VO, which solves camera pose from monocular video based on sparse keypoints. To estimate the scale-consistent relative pose, we present a novel photometric-sensitive depth uncertainty model that accounts for the depth uncertainty introduced by limitations in the photometric error constraint. We also introduce an uncertainty-aware scale recovery strategy that incorporates depth uncertainty for reliable scale alignment. Additionally, we propose a novel difference attention mechanism to construct a point filter that effectively filters out less distinctive points, ensuring high-quality matches for more accurate and efficient pose estimation in the proposed system. Experimental results on the KITTI dataset and Oxford Robotcar dataset demonstrate that our system can predict scale-consistent trajectories from monocular videos and achieve state-of-the-art performance among similar methods. Meanwhile, the depth network within our system achieves competitive depth estimation performance on KITTI depth benchmark.