FastEdit: fast text-guided single-image editing via semantic-aware diffusion fine-tuning

Zhi Chen, Ziwen Zhao, Yadan Luo, Yan Li, Xiaohui Tao, Zi Huang · Pattern Recognition · 2025

• Semantic-aware diffusion: edits guided by image-text discrepancy for precise, controllable changes. • Fast and lightweight: 50-iteration fine-tune with LoRA ( 0.37% params), 17 s per image. • Better alignment without sacrificing fidelity: higher CLIP on TEdBench with competitive LPIPS. • Broadly applicable: runs on SD v1.4 and ports to SDXL (via IP-Adapter) with consistent gains. Text-guided single-image editing has emerged as a promising solution to precisely alter an input image based on the target texts, such as making a standing dog appear seated or a bird spreading its wings. While effective, conventional approaches require a two-step process, including fine-tuning the target text embedding for over 1K iterations and the generative model for another 1.5K iterations. Although it ensures that the resulting image closely aligns with both the input image and the target text, this process often requires 7 minutes per image, posing a challenge for practical application due to its time-intensive nature. To address this bottleneck, we introduce FastEdit, a fast text-guided single-image editing method with semantic-aware diffusion fine-tuning, accelerating the editing process to just 17 seconds. FastEdit streamlines the generative model’s fine-tuning phase, reducing it from 1.5K to 50 iterations. Specifically, we perform diffusion fine-tuning on certain time steps determined by the semantic discrepancy between the input image and target text. We conduct extensive experiments to validate the editing performance of our approach and show promising editing capabilities, including content addition, style transfer, background replacement, and posture manipulation, etc. Our code and more edited images are available at https://fastedit-sd.github.io .

Read the paper · More papers on PaperTik