Cosine Annealing Weights in Knowledge Distillation
Chenglong Wang, Jinhua Wang, Jinhui Xie · 2025
Most existing distillation methods ignore the relationship between the weights of different types of losses in the loss function and fix them as hyperparameters. The choice of weighting parameters has an important impact on model performance, but their default values and the way they are adjusted are not clearly defined. Fixing the weights of the components of the loss function is usually suboptimal for continuously learning student models. In this paper, we innovatively introduce the cosine annealing algorithm into loss weight regulation, specifically, defining the weight ratio of the teacher distillation loss to the student supervision loss to vary as a cosine function with the training cycle. We mimic the human learning process by fixing the weights of the loss function in the first stage to allow the student model to fully learn the output of the teacher model. In the second stage, the L student weights are gradually enhanced to promote practice autonomy. In the third stage, the L distill weights are re-enhanced to realize experience re-refinement and prevent overfitting. This mechanism enables the student model to avoid over-reliance on the teacher's experience while preventing cognitive bias due to local optimality during practice. Experiments in CIFAR-100 demonstrate that this staged learning strategy for the brain significantly outperforms the static weight distillation approach.