High-fidelity face swapping via conditional diffusion model
Yang Zheng, Hongjiang Xian, Xia Yuan, Hongjiang Ma, Jing Hu · 2025
Face swapping has seen rapid progress with the rise of generative models, aiming to seamlessly integrate a source face’s identity with a target face’s expression, pose, and illumination. Although diffusion models exhibit strong generative capabilities, they still face challenges in preserving source identity and achieving natural background integration. In this paper, we propose a high-fidelity face swapping method based on conditional denoising diffusion probabilistic models. A face embedder extracts identity features from the source face and attribute features from the target face, which are integrated via a cross-attention mechanism to guide the generation process. Additionally, a face-background fusion module progressively refines low-frequency details during denoising, enhancing visual coherence. Experiments demonstrate that our method outperforms existing approaches in identity preservation and visual realism.