LanYUAN, a GPT large model using Curriculum Learning and Sparse Attention

Gonghai Zhou, Yuhong Zhang, Rizhen Hu, Yang Zhang · 2023

In 2021, the Inspur AI Research Institute introduced the AI Megatron Model Yuan-1.0, a massive Chinese language AI model containing 245.7 billion parameters. This model surpassed OpenAI's GPT-3, making it the world's largest Chinese NLP model. Although the model was pre-trained using Nvidia's Megatron framework with model parallelism, data parallelism, and pipelining optimizations, there is still room for improvement in terms of training time, cost, and convergence. To achieve better performance, this paper investigates the impacts of batch size and learning rate on model training time and accuracy to balance model performance. We replaced the pipelining optimization with the more efficient DeepSpeed framework, and combined DeepSpeed's ZeRO-based data parallelism with Nvidia's Megatron-LM model parallelism to achieve higher performance on Nvidia GPU clusters with high-bandwidth interconnects. Additionally, we used a curriculum learning-based method and four types of sparse attention as a new optimization approaches. The results showed that the training time was reduced by 20% and the throughput increased by 20% compared to the 47 billion parameters Yuan-1.0 model. Approximately, the optimized model achieved performance improvement in downstream tasks with the same training data.

Read the paper · More papers on PaperTik