RSTNet: Recurrent Spatial-Temporal Networks for Estimating Depth and Ego-Motion
Tuo Feng, Dongbing Gu · IEEE Transactions on Emerging Topics in Computational Intelligence · 2024
Depth map and ego-motion estimations from monocular consecutive images are challenging to unsupervised learning Visual Odometry (VO) approaches. This paper proposes a novel VO architecture: Recurrent Spatial-Temporal Network (RSTNet), which can estimate the depth map and ego-motion from monocular consecutive images. The main contributions in this paper include a novel RST-encoder layer and its corresponding RST-decoder layer, which can preserve and recover spatial and temporal features from inputs. Our RSTNet extracts appearance features from input images, and extracts structure and temporal features from intermediate results for ego-motion estimation. Our RSTNet also includes a pre-trained network to detect dynamic objects from the difference between full and rigid optical flows. A novel auto-mask scheme is designed in the loss function to deal with some challenging scenes. Our evaluation results on the KITTI odometry benchmark show our RSTNet outperforms some of the existing unsupervised learning approaches.