Cross-Vendor GPU Programming: Extending CUDA Beyond NVIDIA

Manos Pavlidakis, Chris Kitching, Nicholas Tomlinson, Michael Søndergaard · 2025

The dominance of NVIDIA's CUDA platform in GPU programming has revolutionized fields such as machine learning (ML), scientific simulations, and computational biology.However, its exclusivity to NVIDIA GPUs poses significant challenges, including vendor lockin, higher costs, and reduced hardware flexibility.Cross-platform solutions like HIP and SYCL often require extensive code rewrites, and while they provide tools to transform CUDA code into their APIs to reduce developer effort, these tools are frequently incomplete and fail to ensure seamless compatibility.Furthermore, the absence of a formal CUDA specification worsens these issues, creating the "CUDA dialect problem", where NVIDIA's nvcc compiler behavior serves as the de facto standard, leading to incompatibilities with alternative compilers.We present SCALE, a platform that enables seamless execution of CUDA applications on AMD GPUs without requiring any code modifications.Leveraging a "dual" compiler approach-a clang-based compiler with language extensions and an nvcc mode enabled in SCALE's clang compiler for existing CUDA codebases-SCALE effectively resolves the CUDA dialect problem.SCALE re-implements CUDA's runtime and driver APIs while mapping NVIDIA's compute capabilities to AMD architectures, ensuring native-level performance.Our evaluation demonstrates SCALE 's capability to support a wide range of real-world frameworks implemented entirely in CUDA.SCALE provides broad compatibility across four AMD GPU microarchitectures and thirteen popular frameworks by supporting the majority of the CUDA API, reflecting 60% coverage.Finally, SCALE delivers minimal performance overhead compared to ROCm, enabling a unified codebase for both NVIDIA and AMD GPUs and greatly simplifying cross-platform development.

Read the paper · More papers on PaperTik