SwapDiffusion: Flexible Swapping Disentangled Content‐Style Embeddings in P+ $\mathcal{P}+$ Space for Diffusion Models

Yongxing He, Zejian Li, Wei Li, Xinlong Zhang, Jia Wei, Yongchuan Tang · IET Computer Vision · 2025

ABSTRACT This paper introduces SwapDiffusion, a novel framework for content‐style disentanglement in diffusion‐based image generation. We advance the understanding of the extended textual conditioning () space in SDXL by identifying the and transformer block layers as primarily responsible for content and style, respectively. Building on this insight, we introduce a novel q‐transformer architecture. It features a block‐diagonal matrix masked self‐attention layer that effectively isolates content and style embeddings by reducing inter‐query interference. This design not only enhances disentanglement but also improves training efficiency. Crucially, the learnt image embeddings align well with textual ones, enabling flexible content and style control via images, text or their combinations. SwapDiffusion supports diverse applications such as style transfer (image‐ or text‐driven), image variation, stylised text‐to‐image generation and multimodal‐prompted image synthesis. Experimental results demonstrate that by aligning learnt image embeddings with the U‐Net's pre‐identified functional layers for content and style, SwapDiffusion achieves superior content‐style separation and image quality while offering greater adaptability than existing approaches. The implementation code and pre‐trained models will be released at https://github.com/lioo717/SwapDiffusion .

Read the paper · More papers on PaperTik