Meta-Learned Dynamic Distillation for Automated Hyperparameter Optimization in Machine Learning Systems

Yulin Zhou · Journal of Engineering Systems and Applications · 2025

We propose a meta-learned dynamic distillation framework for automated hyperparameter optimization in machine learning systems, which adaptively adjusts knowledge transfer intensity between teacher and student models during training. Traditional distillation methods rely on static intensity schedules or manual tuning, often leading to suboptimal performance when faced with dataset shifts or model uncertainty. The proposed method addresses this limitation by formulating distillation intensity as a learnable hyperparameter, optimized in real-time through a bilevel optimization scheme. Our framework integrates three key components: an Adaptive Distillation Controller (ADC) that meta-learns intensity adjustments based on gradient dynamics and validation loss, a bilevel optimization engine that jointly minimizes student loss and intensity regularization, and a curriculum-aware memory buffer that stabilizes training through historical trajectory analysis. The ADC employs a recurrent neural network to dynamically modulate intensity, while the bilevel optimizer ensures efficient meta-gradient computation without unrolled computational graphs. Furthermore, the memory buffer captures gradient variance patterns to inform intensity adjustments, enabling robust adaptation to evolving training conditions. Experiments demonstrate that our method outperforms conventional static distillation approaches across multiple benchmarks, achieving superior model performance with reduced manual intervention. The unified treatment of hyperparameter optimization and knowledge distillation not only eliminates the need for handcrafted schedules but also provides a principled way to balance task-specific learning and teacher guidance. This work advances the state of automated machine learning by introducing a scalable, adaptive solution for model tuning.

Read the paper · More papers on PaperTik