VD-Matcher: A Very Deep Local Feature Matcher With Weight Recycling and Keypoint Detection

Kun Dai, Zilong Zhou, Zhiqiang Jiang, Qihao Sun, Tao Xie, Hongbo Gao, Tao An, Ruifeng Li, Lijun Zhao · IEEE Transactions on Circuits and Systems for Video Technology · 2025

Establishing local feature matches between image pairs serves as a fundamental component of plentiful vision tasks, such as visual localization and Structure from Motion (SfM). Recently, detector-free techniques equipped with the transformer have exhibited exceptional performance. Theoretically, optimizing the transformer architecture and stacking more transformer blocks could emphasize crucial features and filter out extraneous information by progressively narrowing the effective perception regions of the network within images, thereby enhancing the matching performance. Nevertheless, this paradigm results in a linear escalation of model size with respect to the number of blocks. In this study, we introduce VD-Matcher to address this issue. A principal innovation of VD-Matcher is the utilization of a weight recycling technique (WRT) that enables partial weights to be reutilized across successive transformer blocks, along with specific transformations designed to sufficiently enhance feature representations. This approach enables VD-Matcher to construct a deep transformer architecture for accurate local feature matching while maintaining a manageable parameter size. Furthermore, we propose a lightweight multi-scale keypoint detection module that captures representative keypoints to replace all keypoints for compact global information aggregation intra-/inter- images, which reduces the computational overhead induced by excessively deep transformer layers while alleviating redundant information propagation to a certain extent. Extensive experiments verify that VD-Matcher exceeds state-of-the-art algorithms on multiple benchmarks while maintaining less parameters. The source code is available at https://github.com/mooncake199809/VD-Matcher.

Read the paper · More papers on PaperTik