Self-distilled masked autoencoders for medical images
Taeyoung Yoon, Daesung Kang · Engineering Applications of Artificial Intelligence · 2025
Self-supervised learning (SSL) is a transformative paradigm in deep learning, reducing reliance on extensive labeled datasets. As a prominent SSL framework, the masked autoencoders (MAE) leverages masked image modeling to learn robust representations. However, the shallow layers of the MAE encoder struggle to capture sufficient contextual information, which may limit the model's overall performance. To address this limitation, we introduce self-distilled masked autoencoders (SD-MAE), a novel framework that embeds a self-distillation mechanism into the MAE encoder to harness transformer architectures’ inherent structural priors. Inspired by DeepCluster , a self-supervised clustering-based approach that uses alternating clustering and refinement strategy, SD-MAE iteratively refines shallow-layer representations by aligning them with pseudo-labels from deeper layers. Specifically, it first generates pseudo-labels from deep-layer class token features with effectively encoded structural priors, and then propagates this knowledge to refine the shallow-layer features via Kullback-Leibler (KL) divergence minimization. We evaluated the efficacy of SD-MAE on two challenging medical image classification tasks: multi-labeled pediatric thoracic disease classification using chest X-rays and multi-class classification using optical coherence tomography (OCT) images. Experimental results demonstrated that SD-MAE consistently outperformed baseline models, including residual network-34 (ResNet-34), vision transformer-small (ViT-S), and the standard MAE, across various evaluation metrics. Notably, SD-MAE achieved an area under the curve (AUC) of 0.757 for pediatric chest X-ray classification and an AUC of 0.995 for OCT classification, accompanied by significant improvements in sensitivity, precision, and F1-score.