A Lightweight Mamba Backbone for Diffusion Models
Tao Zhou, Yunfan Yu, Zihan Gao · 2024
Recent advances introduced the Visual Mamba (Vim) architecture, featuring bidirectional Mamba blocks, which has demonstrated exceptional performance across various computer vision tasks. This paper extends the application of Vim to diffusion models. However, Vim's sequential processing design overlooks spatial continuity and local structure in images. To address this, we propose AMamba, a novel diffusion model backbone that combines the computational efficiency of Vim with the powerful attention mechanisms of U- ViT, aiming to overcome U-ViT's high computational complexity and parameter count while enhancing spatial continuity modeling. By replacing U-ViT's MLPs with Vim blocks, AMamba efficiently captures sequential dependencies inherent in diffusion processes. Furthermore, time-step and conditional information are embedded directly into the input sequence, enhancing information flow during the diffusion process. The architecture also incorporates long skip connections across layers, which facilitate efficient information propagation and contribute to faster convergence. Experimental results show that AMamba achieves significantly lower FID scores on the CelebA64 dataset compared to U-ViT and Vim, while requiring fewer training iterations and significantly reducing the number of model parameters.