Parallel Processing for Distributed Machine Learning: A Taxonomy of Techniques and Associated Security Risks

Abdulfatah Bahbouh, Ishfaq Ahmad, Hansheng Lei, Saif ul Islam · 2025

As the scale of both data and models surpasses the capabilities of edge devices, the utilization of parallel and distributed systems has become increasingly crucial. This paper overviews the landscape of Deep Learning (DL) with parallel processing, its evolution, and integration with cloud computing platforms. This review aims to provide researchers and practitioners with a thorough understanding of the current state and future directions of cloud-based parallelization in DL. It provides a taxonomic classification of current parallelization methods, including data parallelism, layer parallelism, task parallelism, model parallelism, pipeline parallelism, subgraph parallelism, and hybrid approaches. The details include advanced implementations such as tensor parallelism in Megatron-LM, as well as additional frameworks such as GSPMD, and DAPPLE. The paper delves into the intricacies of these methods, discussing their strengths, limitations, and applications. Furthermore, it examines the distribution of large models across edge devices and cloud infrastructure, addressing the inherent privacy risks and computational challenges. The paper also considers privacy-preserving techniques such as federated learning (FL) and split learning (SL), which enable collaborative model training while maintaining the locality of personally identifiable information (PII). The research highlights specific implementations like FedSL, which combines federated and SL to handle non-IID data distributions. By synthesizing recent advancements and challenges, the paper offers insights into the development of scalable, efficient, and privacy-preserving solutions for large-scale DL systems in the era of big data and stringent data protection regulatory landscape.

Read the paper · More papers on PaperTik