OptimusNIC: Offloading Optimizer State to SmartNICs for Efficient Large-Scale AI Training

Achref Rebai, Marco Canini · 2025

LLM training is a demanding workload that requires careful coordination among hardware components, ensuring high GPU utilization, rapid data transfer, and minimal memory overhead. This synergy leads to an efficient AI training system. Scaling LLMs encounters memory constraint challenges, particularly the optimizer state, whose size in bytes scales with a factor of 12× the number of model parameters. In this paper, we explore the impact of offloading the optimizer state and parameter update operation to a SmartNIC. OptimusNIC reduces communication overhead between GPUs, minimizes GPU computation (by handling the optimizer step externally), and significantly decreases memory requirements. These improvements allow GPUs to accommodate additional model layers and efficiently train and fine-tune models with minimal resources. In addition to examining the impact of OptimusNIC, we evaluate DeepSpeed's ZeRO-Infinity framework, identifying key limitations associated with using the CPU as the offloading target and we demonstrate why OptimusNIC offers a more efficient alternative. Furthermore, we analyze the usability of one of the target platforms for OptimusNIC: the NVIDIA BlueField-2 DPU.

Read the paper · More papers on PaperTik