Efficient Multi-GPU Programming in Python: Reducing Synchronization and Access Overheads

Lena Oden, Klaus Nölp · 2025

Python has become increasingly significant in domains such as data science, machine learning, scientific computing, and parallel programming. The libraries CuPy and Numba enable the development of parallel GPU code, while mpi4py and CuPy's NCCL backend enable distributed computing across multiple GPUs. Despite its versatility, Python is often criticized for its performance limitations. Although pre-compilation and just-in-time compilation can minimize interpreter overhead, multi-GPU applications in Python often encounter significant performance bottlenecks due to the synchronization requirements between GPU kernels and communication libraries. In this work, we present a detailed performance analysis of multi-GPU programming in Python using CuPy, Numba, NCCL and mpi4py. We identify excessive synchronization and costly array conversions as key sources of overhead and demonstrate that view-based data access can significantly improve performance. Furthermore, we show that using NCCL with asynchronous CUDA streams enables better overlap of computation and communication, mitigating interpreter-induced delays. Our evaluation includes both microbenchmarks and a multi-GPU implementation of the CloverLeaf mini-application. Results show that, with careful optimization, Python implementations can reach up to 90 % of the performance of equivalent C-CUDA codes. These findings highlight practical strategies for minimizing Python-specific overheads in multi-GPU scenarios and provide guidance for building efficient Python applications on modern GPU clusters.

Read the paper · More papers on PaperTik