DHT: Deformable Hybrid Transformer for Aerial Image Segmentation
Yan Zhang, Xiyuan Gao, Qingyan Duan, Lin Yuan, Xinbo Gao · IEEE Geoscience and Remote Sensing Letters · 2022
Due to the strong ability to model global information, the transformer-based methods have shown remarkable improvements in image segmentation tasks. However, the self-attention mechanism in the transformer is computationally expensive and relies on pre-trained parameters. Moreover, the transformer method is weak in modeling local information, which is unfavorable for accurately segmenting objects from high-resolution aerial images. To this end, an efficient deformable orientational self-attention (DoA) is proposed to simultaneously extract the global information and the local information. Besides, for parameter efficiency, we design a depthwise channel self-attention (DcA) to model the contextual information among channels. Combining with the DoA and DcA, we propose the deformable hybrid transformer (DHT) to perform high-quality object segmentation on aerial images. Experiments on ISPRS Potsdam dataset and WHU building dataset illustrate that the proposed DHT can not only achieve state-of-the-art (SOTA) results but also markedly reduce the dependence of the transformer on pre-trained parameters.