Multi-Scale Information Can Do More: Attention Aggregation Mechanism for Image Compression
Bo Li, Yongjun Li, Yao Li, Jincheng Luo, Chaoyue Li · 2023
The semantic information obtained from large-scale computation in image compression is not practical. To solve this problem, we propose an Attention Aggregation Mechanism (AAM) for learning-based image compression, which is able to aggregate attention map from multiple scales and facilitate information embedding. Based on AAM and Swin Transformer, we design a Multi-scale Self-Attention Image Compression (MSAIC) approach as shown in Figure 1, which is using Multi-scale Swin Transformer blocks (MSTB) and convolutional layers stacks in the down-sampling encoder and up-sampling decoder to the best embed information. The MSTB is made up of convolution layers and three shifted window multi-head self-attention (SW-MSA) blocks, each of which has a different number of heads. In this way, a single MSTB block can aggregate three scales of attention map from tiny, middle and large patches. It can also function as a plug-and-play component to enhance CNN-based models. Meanwhile, we use the channel-wise autoregressive entropy model for the efficient and accurate entropy probability estimation. Extensive experiments show MSAIC is effective and attain the state-of-the-art (SOTA) compression performance, also outperform the SOTA method at the high bit rate.