One-stage detection for unsupervised domain adaptation with efficient multi-scale attention and confidence-augmented combination

Nan Xiang, Qianxi Liu, Yaoyao Jiang · Journal of Electronic Imaging · 2024

Unsupervised domain adaptation for object detection leverages a labeled domain to learn an object detector generalizing to a different domain free of annotations. We propose efficient multi-scale attention, confidence mixing, augmentation, and combination (ECAC), an adaptive object detector learning method based on a region-level confidence sample mixing strategy. Compared with the current methods, our approach crops high-confidence detection regions from both the source and target domains, augments them, and combines them to generate composite samples. In addition, consistency loss is utilized to solve the domain adaptation problem. Furthermore, we introduce the efficient multi-scale attention (EMA) into the detector. To retain channel information and reduce computational overhead, EMA attention restructures part of the channels into the batch dimension and groups the channel dimension into multiple sub-features, ensuring spatial semantic features are evenly distributed within each feature group. EMA employs a shared 1×1 convolution branch from the CA attention module, along with a parallel 3×3 convolution kernel to aggregate multi-scale spatial structure information. This approach effectively enhances the model’s focus on region-level features by integrating local and global information with multi-scale parallel sub-networks and cross-spatial learning. For pseudo-label filtering, we progressively transition from a loose to a stricter confidence threshold. Initially, this allows more pseudo-labels, facilitating the detector’s learning of target domain representations. As training progresses, stricter thresholds are applied to select more reliable pseudo-labels, gradually filtering out inaccurate pseudo-detections. Our extensive experiments on three datasets demonstrate that ECAC achieves state-of-the-art performance on two of them. On the third dataset, our method improves the mean average precision by over 2% compared with the latest methods.

Read the paper · More papers on PaperTik