Large-Scale AI Infra Reliability: Challenges, Strategies, and Llama 3 Training Experience
Xun Jiao, Abhinav Pandey, Karthik Pattabiraman, Fred Lin · 2025
As AI infrastructure grows in complexity and scale, particularly with the training of large language models (LLMs) that require large-scale GPU clusters, ensuring their reliability becomes crucial. In large-scale training jobs, a single GPU failure can potentially disrupt the entire process, impacting tens of thousands of interconnected GPUs. The growing complexity of the hardware further complicates failure attribution and mitigation, exacerbating the effects of hardware failure. This paper begins with a literature review of the current challenges in AI infrastructure and recent industrial efforts to enhance the reliability and fault tolerance of large-scale AI infrastructure, from companies such as Meta, Microsoft, ByteDance, and Alibaba. The review underscores the necessity for ensuring the reliability of large-scale AI infrastructure. We then present our practical experience in training Llama 3 on a 16K GPU cluster, analyzing the types of hardware failures encountered and their root causes. Finally, we discuss three strategies to improve the reliability of large-scale AI infrastructure, informed by our literature review and practical experience.