Sorting Large Datasets with Heterogeneous CPU/GPU Architectures

Michael Gowanlock, Ben Karsin · 2018

We examine heterogeneous sorting for input data that exceeds GPU global memory capacity. Applications that require significant communication between the host and GPU often need to obviate communication overheads to achieve performance gains over parallel CPU-only algorithms. We advance several optimizations to reduce the host-GPU communication bottleneck, and find that host-side bottlenecks also need to be mitigated to fully exploit heterogeneous architectures. We demonstrate this by comparing our work to end-to-end response time calculations from the literature. Our approaches mitigate several heterogeneous sorting bottlenecks, as demonstrated on single- and dual-GPU platforms. We achieve speedups up to 3.47× over the parallel reference implementation on the CPU. The current path to exascale requires heterogeneous architectures. As such, our work encourages future research in this direction for heterogeneous sorting in the multi-GPU NVLink era.

Read the paper · More papers on PaperTik