Axial-shunted Spatial-temporal Conversation for Change Detection

Yuchao Feng, Mengjie Qin, Jiawei Jiang, Jintao Lai, Jianwei Zheng · ACM Transactions on Multimedia Computing Communications and Applications · 2025

Benefitting from the maturing of intelligence techniques and advanced sensors, recent years have witnessed the full flourishing of change detection (CD) on multi-temporal remote sensing images. However, extraneous interference caused by normal temporal evolution and the extreme sparsity of spatial changes still plague the detection accuracy. To counteract this dilemma, a lightweight axial-shunted spatial-temporal conversation network (ASCNet) is proposed, which models the intrinsic representations in dually augmented images with a parallel treatment of convolutions and attentions. Specifically, for the features of weakly augmented bi-temporal image pairs from Siamese CNN, a roundtable attention-based and intra-scale axial-shunted interaction, with linear complexity, is presented. By splitting horizontally or vertically into multiple chunks and then performing axial-squeeze operation, axial-shunted scheme can achieve fine-grained attention while maintaining linear complexity. Moreover, roundtable attention pursues efficient bi-temporal modeling by incorporating both self-attention and cross-attention in a single attentional computation, while imposing change guiding and difference gating for focusing on changes. Simultaneously, a video transformer is introduced for the modeling of strongly augmented sequences, followed by an inter-scale spatial-temporal alignment to recalibrate the feature responses. ASCNet demonstrates state-of-the-art performance on four publicly available CD datasets while maintaining superior computational efficiency. The source code is available at https://github.com/fengyuchao97/ASCNet .

Read the paper · More papers on PaperTik