S-VFNet: a siamese object detection method based on salient features and visual features
Huanhuan Wang, Lisheng Jin, Xinyu Sun, Yang He · Measurement Science and Technology · 2025
Abstract Object detection is a critical application in the field of visual sensing measurements. Despite significant advancements in detection algorithms over recent years, challenges remain in complex traffic scenarios, particularly concerning inadequate feature representation, especially in environments characterized by high target density, severe edge occlusion, and blurred foreground–background distinctions. To address these challenges, this paper proposes a novel siamese object detection network guided by image salient features (S-VFNet). First, an image saliency map is extracted from the original RGB image to generate multi-modal input data. A siamese backbone network based on WTConv is devised, processing the original visual features and the saliency features as separate input streams, thereby facilitating cross-modal feature complementarity. Next, a multi-resolution feature aggregation module, a triple feature encoding module, and a cross-dimensional attention (CDA) module are constructed. The multi-resolution feature aggregator builds a sequential representation of 2D features at different resolutions and uses 3D convolution for scale-aware feature extraction, enhancing multi-level information fusion. The triple feature encoder scales the features across three dimensions to capture local details in densely packed object scenes. Furthermore, a CDA module is introduced to refine spatial localization by capturing channel context correlation. Finally, experiments on the VRU and KITTI datasets demonstrate that the mAP of S-VFNet is 95.6% and 96.2%, showing that the detection performance of S-VFNet outperforms most advanced methods.