DMCNet: Toward Lightweight Volumetric Stereo Matching

James Okae, Huabiao Qin · IEEE Sensors Journal · 2024

Recent methods in stereo matching have steadily improved disparity prediction accuracy using volumetric deep convolutional neural network (CNN) architectures. However, this performance gain is achieved with a large memory footprint, making models from existing volumetric stereo networks unable to run on memory-constrained embedded vision devices. This challenge raises an important research question that inspires this work: Can we achieve a lightweight volumetric deep CNN architecture without sacrificing accuracy? To answer this question, we propose a discriminative multiscale context network (DMCNet), a memory-efficient backbone for volumetric stereo matching. Our key insight is to leverage a lightweight multibranch network to extract rich contextual information and a series of downsampling layers to achieve sufficient semantics with a large receptive field. In addition, we design a feature reuse and fusion (FRF) module that combines information from complementary scales to obtain a more robust representation for accurate disparity regression. We show that these network design recipes provide stronger representations with fewer parameters for learning high-quality disparity map prediction. Experiments on stereo benchmarks reveal that the proposed DMCNet achieves better parameter and memory efficiency as well as competitive accuracy compared to many existing state-of-the-art methods.

Read the paper · More papers on PaperTik