ACTA-AOD: Asymmetric Convolution–Triple Attention Network for Non-Uniform Single-Image Dehazing via Windowed Efficient Multi-Scale Attention

Yuanying Zhang, Fuxing Yu, Yina Suo · Applied Sciences · 2026

Single image dehazing remains a fundamental challenge in computer vision due to the ill-posed nature of the inverse problem and the spatial heterogeneity of real atmospheric haze. Existing convolutional approaches suffer from two structural deficiencies: bounded receptive fields that fail to model large-scale haze gradients, and isotropic kernels insensitive to the directional patterns of atmospheric scattering. This paper proposes ACTA-AOD, a lightweight end-to-end dehazing network that addresses both limitations within a unified framework built upon the AOD-Net K-parameterization. The network integrates two complementary modules: (1) W-EMSAv2, a windowed efficient multi-scale attention module that reduces attention complexity from O(N2C) to O(NM2C/4) while preserving full-spectrum spatial information through pixel-shuffle reconstruction; and (2) the ACTA Fusion module, which combines structural-reparameterization-based asymmetric convolution with cross-dimensional Triple Attention for direction-sensitive local detail recovery at zero inference-time overhead. On the RESIDE benchmark, ACTA-AOD achieves peak signal-to-noise ratio (PSNR) of 26.02 dB and structural similarity index measure (SSIM) of 0.910 on indoor synthetic data, and 26.13 dB/0.910 on outdoor synthetic data, surpassing the AOD-Net baseline by +3.41 dB (indoor) and +3.58 dB (outdoor) in PSNR, and exceeding the strongest learning-based baseline (AECRNet, CVPR 2021) by +1.17 dB (indoor) and +1.75 dB (outdoor). The model processes images at 81 frames per second on a single GPU. Ablation studies and stratified robustness evaluation across five haze density levels confirm the complementary, synergistic contribution of each module.

Read the paper · More papers on PaperTik