U-Sketch: Bridging the Gap for Efficient and Realistic Sketch-to-Image Synthesis

Ilias Mitsouras, Eleftherios Tsonis, Paraskevi K. Tzouveli, Athanasios S. Voulodimos · IEEE Access · 2026

Diffusion models have demonstrated remarkable performance in text-to-image synthesis, producing realistic, high-resolution images that faithfully adhere to the corresponding text-prompts. However, they still fall behind in sketch-to-image synthesis tasks, where, in addition to text-prompts, the spatial layout of the generated images has to closely follow the outlines of certain reference sketches. Current adapter-like networks, such as ControlNet and T2I-Adapters, are highly computationally intensive to train and often struggle to consistently yield realistic results. Conversely, prior Multi-Layer Perceptron (MLP)-based latent edge predictors operate on a per-pixel basis, failing to capture holistic spatial layouts and resulting in severe time inefficiency due to the numerous denoising iterations required. To address these limitations, we introduce U-Sketch, an efficient framework featuring a U-Net latent edge predictor. During inference, this predictor extracts intermediate feature maps from the pre-trained diffusion model’s denoising network to guide the generation process. This allows U-Sketch to accurately capture both local and global spatial pixel correlations without requiring base model retraining. The proposed framework also incorporates a sketch simplification network that allows users to preprocess and simplify input sketches for enhanced outputs, as well as a noise initialization strategy that aligns the intrinsic spatial layout of the noise with the desired structure, thereby facilitating the denoising process. Extensive quantitative and qualitative evaluations demonstrate our key contributions: U-Sketch consistently generates more realistic images with superior edge fidelity compared to existing baselines. Notably, it drastically reduces the required denoising steps by 80% compared to MLP-based methods and requires minimal training resources, allowing users to easily fine-tune the framework for their distinct handwriting styles. To the best of our knowledge, U-Sketch is the first lightweight, U-Net-based framework that successfully resolves the trade-off between computational efficiency and high-fidelity spatial alignment in the sketch-guided text-to-image synthesis task.

Read the paper · More papers on PaperTik