SATURN: Autoregressive Image Generation Guided by Scene Graphs

Thanh-Nhan Vo, Trong-Thuan Nguyen, Tam Nguyen, Minh–Triet Tran · 2025

State-of-the-art text-to-image models can generate photorealistic images, yet they often struggle to accurately capture the layout and object relationships implied by complex prompts. Scene graphs provide the missing structural signal. However, prior graph-guided generators have typically relied on heavy GAN or diffusion architectures, which lag behind modern autoregressive pipelines in terms of speed and fidelity. In this work, we introduce SATURN (Structured Arrangement of Triplets for Unified Rendering Networks), a lightweight, drop-in extension to VAR-CLIP that translates a scene graph into a salience-ordered token sequence, enabling a frozen CLIP–VQ-VAE backbone to parse the graph while fine-tuning only the VAR transformer. Empirically, on the Visual Genome dataset, SATURN reduces the FID score from 56.45% to 21.62% and raises the Inception Score from 16.03% to 24.78%, achieving gains that outperform SG2IM and SGDiff without requiring additional modules or multistage training. Notably, our qualitative results further confirm improvements in object count fidelity and spatial relation accuracy.

Read the paper · More papers on PaperTik