SSDC: A Scalable Sparse Differential Checkpoint for Large-scale Deep Recommendation Models

Lingrui Xiang, Xiaofen Lu, Rui Zhang, Zheng Hang Hu · 2024

Today deep recommendation models have become increasingly large, with parameter sizes reaching hundreds of GB or even TB scale. As a result, it requires large-scale cluster computing resources to train such models. However, large-scale computing clusters tend to experience frequent failures during runtime, so fault-tolerance mechanisms such as checkpointing and restart are widely used in model training. Traditional checkpointing techniques periodically save all parameters of model, resulting in significant overhead. To address this issue, we propose an improved partial checkpointing mechanism for recommendation models named SSDC. SSDC uses an adaptive threshold strategy to reduce expensive operations when saving checkpoints, thereby having good scalability. Furthermore, SSDC saves the differential value of the model parameters, making it feasible to sparsify the otherwise dense embedding tables, thus reducing the bandwidth and time overhead to reconstruct checkpoints. Our evaluations show that compared to state-of-the-art methods, SSDC greatly reduces the time overhead of saving and reconstructing checkpoints, while achieving comparable training accuracy.

Read the paper · More papers on PaperTik