SpIRL: Spatially-aware image representation learning under the supervision of relative position descriptors

Logan Servant, Michaël Clément, Laurent Wendling, Camille Kurtz · Pattern Recognition · 2025

Extracting good visual representations from image contents is essential for solving many computer vision problems (e.g. image retrieval , object detection, classification). In this context, state-of-the-art approaches are mainly based on learning a representation using a neural network optimized for a given task. The encoders optimized in this way can then be deployed as backbones for various downstream tasks. When the latter involves reasoning about spatial information from the image content (e.g. retrieve similar structured scenes or compare spatial configurations), this may be suboptimal since models like convolutional neural networks struggle to reason about the relative position of objects in images. Previous studies on building hand-crafted spatial representations , thanks to Relative Position Descriptors (RPD), showed they were powerful to discriminate spatial relations between crisp objects, but such spatial descriptors have rarely been integrated into deep neural networks . We propose in this article different strategies embedded in a common framework called SpIRL (SPatially-aware Image Representation Learning) to guide the optimization of encoders to make them learn more spatial information, under the supervision of an RPD and with the help of a novel dataset (44k images) that does not induce learning semantic information. By using these strategies, we aim to help encoders build more spatially-aware representations. Our experimental results showcase that encoders trained under the SpIRL framework can capture accurate information about the spatial configurations of objects in images on two selected downstream tasks and public datasets.

Read the paper · More papers on PaperTik