Self-Supervised Photographic Image Layout Representation Learning
Zhaoran Zhao, Peng Lu, Xujun Peng, Wenhao Guo · IEEE Transactions on Multimedia · 2025
Image layout representation learning, which converts layouts into compact vectors, is essential for tasks such as image retrieval, editing, and generation. However, existing methods—especially those applied to photographic images—face several challenges: supervised methods rely on expensive labeled datasets, weakly-supervised methods struggle with generalization, and self-supervised methods are limited in handling the diversity of photographic layouts. To address these issues, we propose a novel heterogeneous layout graph that efficiently captures the layout information in images. The vertices of this graph represent the compositional primitives of the image, capturing their attributes, while the edges encode the relationships between these primitives. We also design effective pretext tasks to guide a layout encoder-decoder in self-supervised training, ultimately generating the layout graph embedding vector. Additionally, we introduce a new layout evaluation dataset—LODB—which features a richer variety of layout categories, significantly better label quality than existing datasets, and a more balanced distribution of semantic scenes across layout categories, providing a comprehensive benchmark for evaluation. Experiments on the LODB dataset demonstrate that our method outperforms existing approaches in representing photographic image layouts.