Distributed Learning for Large-Scale Models at Edge With Privacy Protection
Yuan Yuan, Shuzhen Chen, Dongxiao Yu, Zengrui Zhao, Yifei Zou, Lizhen Cui, Xiuzhen Cheng · IEEE Transactions on Computers · 2024
Big data and strong computing power have promoted artificial intelligence to the era of big models. In particular, ChatGPT’s debut heralded the vigorous development of large models. It is an urgent problem to train large models with trillion-level parameters efficiently. Traditional single-machine training stores all data and model parameters in memory. However, due to the limitation of memory and communication resources, when the amount of data or model parameters increases, the problem of memory shortage and communication blocking often occurs. Therefore, distributed training is the most effective ways to solve the above problems and improve training efficiency. In this paper, we propose the algorithmDL-DP, which can achieve an asymptotically optimal convergence rate$O(1/{\sqrt{TK\Gamma^*}})$while satisfyingε-differential privacy, whereTis the local epoch number,Kis the global maximum iteration number and$\Gamma^*$is the minimum covering index. In particular, when${\Gamma ^*} = N$, DL-DP achieves a convergence rate of$O(1/{\sqrt{TKN}})$, which is equivalent to the best-known FedAvg approach implemented by training the full model at each client. When${\Gamma ^*} = 1$, DL-DP achieves a convergence rate of$O(1/{\sqrt{TK}})$, which is comparable to OAP that assumes all parameters need to be trained at least once in each iteration. Finally, our algorithm has been demonstrated to converge through extensive experiments.