Layer Dance: Pairwise Layer Training with Adaptive Learning Rates for Transformers

Aakash Singh · 2025

A new method for training decoder-only transformer models for sequence generation tasks is proposed. An Adaptive Learning Rate (ALR) scheduler is introduced, in which learning rates are dynamically adjusted based on gradient and activation norms. Additionally, a Pairwise Layer Training strategy is employed, where only two adjacent decoder layers are trained at a time, and layers are cycled through systematically. Experiments conducted on datasets containing 5,000 and 10,000 samples demonstrate that the combination of ALR and Pairwise Layer Training results in lower training loss and faster convergence compared to traditional training methods. The implementation is publicly available at https://github.com/ aakash-dec7/Adaptive-Learning-Rates.

Read the paper · More papers on PaperTik