D-Net: A Generalised and Optimised Deep Network for Monocular Depth Estimation
Joshua Luke Thompson, Son Lam Phung, Abdesselam Bouzerdoum · IEEE Access · 2021
Depth estimation is an essential component in computer vision systems for achieving 3D scene understanding. Efficient and accurate depth map estimation has numerous applications including self-driving vehicles and virtual reality. This paper presents a new deep network, called D-Net, for depth estimation from a single RGB image. The proposed network is designed as an efficient, accurate and universal model that can adopt a wide range of encoder backbones. Our approach gathers strong global and local contextual features at multiple resolutions and transfers these to high resolutions for clearer depth maps. For the encoder backbone we adopt state-of-the-art models including EfficientNet [1], HRNet[2] and Swin Transformer [3] to obtain densely labelled depth maps. The proposed D-net can be trained end-to-end and is designed to have minimal parameters and a reduced computational complexity. Extensive evaluations on the NYUv2 [4] and KITTI [5] benchmark datasets show that our model is highly accurate across multiple backbones and achieves state-of-the-art performance on both benchmark datasets when combined with the Swin Transformer and HRNets.