Masked image modeling framework with semantic similarity-preserving
Sijia Wang, Huihu Shao, Yun Ge, Yuanyuan Shen · 2025
To reduce the annotation costs caused by the surge in images, Masked Image Modeling (MIM) has been employed to leverage unlabeled data for training, with trained weights then transferred to downstream tasks. However, conventional MIM struggles with semantic inconsistency between predicted and original images. To address this, a Semantic Similarity-Preserving Masked Image Modeling framework (SPM) is proposed. SPM consists of image reconstruction and semantic preservation components. In image reconstruction, the Masked Autoencoder (MAE) predicts masked images. The original and predicted images are then fed into CLIP within the semantic preservation component to predict their category probabilities. To ensure semantic consistency, a novel Semantic Preservation Loss (SPLoss) is introduced to optimize semantic differences. The model is trained by jointly optimizing MSELoss and SPLoss. Experiments show that the SPM framework outperforms the classical MAE framework in retrieval performance on two datasets. Notably, compared to a conventional pre-trained ViT using 80% labeled data, SPM achieves comparable retrieval accuracy with 20% less labeled data, even slightly improving in some cases. This effectively mitigates the high costs of annotated images.