MSAttU-Net: A Water Body Extraction Network for Nanjing That Overcomes Interference in Shaded Areas
Yiheng Xie, Hongyue Zhang, Xiaoping Rui, Heng Tang, Ninglei Ouyang, Yarong Zou · IEEE Geoscience and Remote Sensing Letters · 2025
The extraction of water bodies in densely built-up areas is often surrounded by complex background environments, with shadow regions caused by high-rise buildings in high-resolution remote sensing images being one of the most serious challenges in water body recognition research. Convolutional neural networks (CNNs) have significant research value in capturing spatial structures in images. Compared to other deep learning models, CNNs require fewer parameters and computational resources while effectively extracting local features, achieving satisfactory segmentation performance. However, in the existing high-resolution remote sensing images, the color and shape of shadow regions exhibit high similarity to water bodies. The use of single-scale structures and shared weights severely hinders the CNN model’s ability to accurately distinguish and extract water bodies under strong interference conditions. To address these issues, this letter proposes a lightweight CNN method (MSAttU-Net) to overcome the interference of urban shadow regions in large-scale high-resolution remote sensing images. First, a lightweight multibranch deep convolution module is constructed to expand the model’s receptive field and enhance its feature extraction capability. Second, a parallel attention mechanism module (P-Att) is introduced, consisting of parallel spatial and channel attention mechanisms. These are designed to improve the focus on water boundary and shape features as well as corresponding spectral information, reducing the interference from shadows and other noise in the images. Finally, a$1\times 1$convolution layer is used to generate pixel-level classification results. The proposed method’s performance is compared with that of contemporary CNN networks, including U-Net, Seg-Net, Res-Net, and Dense-Net, popular CNN networks, such as DeeplabV3+, HR-Net, and SegFormer, and the state-of-the-art (SOTA) networks, such as TransUnet, SwinUnet, and DS-TransUnet. The results indicate that the proposed model not only maintains a lightweight and efficient structure but also demonstrates exceptional capability in water body recognition and effectively overcoming shadow region interference.