GPU Simulation Acceleration via Parallelization
Rodrigo Huerta, Antonio M. González · 2025
Simulating modern GPU architectures with increased core counts and recent workloads can be challenging, even on powerful computing platforms. In this paper, we present a simple approach to parallelize Accel-sim with minimal code changes using OpenMP. In addition, we introduce PaSSMA, a novel technique to optimize the OpenMP for-loop scheduler performance in the different simulated workloads. Moreover, our parallelization technique is deterministic, so the simulator provides exactly the same results for single-threaded and multithreaded simulations. When we run the simulator with 16 CPU cores, we achieve an average speed-up of 6.4x and reach 10x in some workloads.