Parallel Computing Data Processing: Frameworks, Implementations, and Case Studies

Prudhvi Naayini · International Journal of Advances in Engineering and Management · 2025

Parallel computing has become fundamental in processing the massive data volumes generated in modern science and industry. This paper presents a comprehensive survey and practical review of parallel computing in data processing, examining key frameworks (MPI, Open MP, CUDA, Hadoop MapReduce, Apache Spark, etc.), implementations, and real- world case studies. We discuss the architectures underlying shared-memory and distributed-memory systems and illustrate how parallelism on CPUs and GPUs is exploited for high- performance computing (HPC) and big data analytics. We review theoretical foundations (including Amdahl’s law for speedup limits) and compare programming models in terms of design, performance, and fault tolerance. A two-column IEEE-style for- mat is used to present critical insights from scholarly literature, with diagrams illustrating parallel architectures and workflows, and tabular summaries of tools and results. Case studies—from scientific HPC simulations to large-scale data analytics—highlight practical parallelism on multicore CPUs, GPUs, and computing clusters. We conclude with a discussion on the convergence of HPC and big data paradigms, the challenges of heterogeneous computing, and emerging trends. This survey serves as a reference for PhD researchers and practitioners seeking to understand frameworks and real-world practices in parallel computing for data processing.

Read the paper · More papers on PaperTik