Dual-Dimensional Self-Attention for Facial Attribute Editing
Zhenxiong Chang · 2024
Currently, methods based on the Enhanced Vision Transformer with Dual-Dimensional Self-Attention effectively capture long-range dependencies in images by computing self-attention mechanisms across multiple dimensions. This approach demonstrates superior performance in tasks such as image classification and object detection by partitioning images into small patches and treating them as sequential data. In this paper, we propose a framework that utilizes Dual-Dimensional Self-Attention as an encoder for image inversion. Unlike the complex model structure of encoder4editing, we employ a single modular component and use a simple network to reconstruct and edit real images. Experimental results indicate that this method excels in high-quality image inversion, with only a slight reduction in editing accuracy. Its reconstruction and editing performance are both slightly better than the complex structure based on encoder4editing.