Research on Large Model Resource Optimization Strategy Through Distributed Management
Jinhai Wang · 2025
This paper mainly discusses how to optimize the use of resources effectively through distributed management in large-scale machine learning model training and reasoning. First, this paper analyzes the challenges that traditional resource management approaches face when dealing with large models, such as compute pressure, storage requirements, and network bottlenecks. Then, a new resource optimization strategy based on distributed architecture is proposed, which utilizes dynamic load balancing, task scheduling and resource sharing technologies to improve resource utilization and system efficiency. In addition, this paper also demonstrates the application effect of the proposed strategy in different scenarios through case analysis and experimental verification, and proves its advantages in reducing cost, shortening training time and improving model performance.