Parallel Spectral Graph Partitioning on CUDA
Orhan Fırat, Alptekin Temizel · 2014
CPU GPU Configuration 1: Step 1.1 thrust::sort_by_key Step 1.2 memory operations (host to device) Step 1.3 kernel executions Step 1.4 memory operations (device to host) Step 1.5 matrix symmetrization Step 1.6 setting nearest neighbors on distance matrix Step 1.7 thrust::sequence Step 2.1 memory operations (host to device) Step 2.2 kernel executions Step 2.3 memory operations (device to host) Step 3.1 kernel executions Step 3.2 memory operations (host to device) Step 3.3 memory operations (device to host) Step 4.1 conversion of eigenvector matrix Step 4.2 thrust::sort Step 4.3 thrust::sequence Step 5.1 normalize matrix rows Step 5.2 cuda-kmeans using block shared mem. opt. Overall Run Time