Bootstrap Your Own Teacher: Online Policy Distillation for Multi-Game Reinforcement Learning
Donal Byrne, Marko Tot, Paul Duckworth, Clément Bonnet, Alexandre Laterre, Thomas D. Barrett · 2025
Training generalist agents capable of performing well across diverse environments is a significant goal of reinforcement learning (RL). Current state-of-the-art methods for multi-game RL rely on offline datasets, and often discard the policy used to gather trajectories despite its potential to provide a rich learning signal. In this paper, we revisit policy distillation (PD) for multi-game RL and introduce a new framework called Bootstrap Your Own Teacher (BYOT) that extends policy distillation to the online-RL setting. BYOT alternates between two phases: (i) game-specific finetuning and (ii) distilling bootstrapped teachers back into a shared multi-game policy. By directly regulating the multi-game learning dynamics in policy space, BYOT balances training without explicit gradient adjustments or reward normalization, whilst being highly parameter efficient. Our framework is empirically validated for both online and offline multi-game learning on the Atari-40 benchmark. BYOT outperforms all prior online Atari-40 multigame agents, achieving an IQM human-normalized-score (HNS) of 152.7 %. Moreover, by adopting state-of-the-art PPO teacher agents—contrasting the widely-used datasets from weaker DQN agents—and policy distillation, we more than triple the leading IQM-HNS in offline settings to 369.5 %, whilst using significantly fewer parameters. Overall, our results emphasise the power of distillation in multi-game settings.