An Empirical Study of Issues in Large Language Model Training Systems
Yanjie Gao, Ruiming Lu, Haoxiang Lin, Yueguo Chen · 2025
Large language models (LLMs) have gained significant traction in recent years, driving advancements in various applications. The training and evaluation of these models depend heavily on specialized LLM training systems, which are deployed across numerous GPUs, partition LLMs, and process large datasets. However, issues in LLM training systems can lead to program crashes or unexpected behavior, reducing development productivity and wasting valuable resources such as GPUs and storage.