Seeing the whole in the parts with self-supervised representation learning
Arthur Aubret, Céline Teulière, Jochen Triesch · Neurocomputing · 2026
Humans learn to recognize categories of objects, even when exposed to little language supervision. Behavioral studies and the successes of self-supervised learning (SSL) models suggest that this learning may hinge on modeling spatial regularities of visual features. However, SSL models rely on geometric image augmentations such as masking portions of an image or aggressively cropping it, which are not known to be performed by the brain. Here, we propose CO-SSL, an alternative to geometric image augmentations to model spatial co-occurrences. CO-SSL aligns local representations (before pooling) with a global image representation. Combined with a neural network endowed with small receptive fields, we show that it outperforms previous methods by up to on ImageNet-1k when not using cropping augmentations. In addition, CO-SSL can be combined with cropping image augmentations to accelerate category learning and increases the robustness to internal corruptions and small adversarial attacks. Overall, our work paves the way towards a new approach for modeling biological learning and developing self-supervised representations in artificial systems.