Hardware Hybrid Computing solutions

B. Stefanizzi · 2008

Parallel HPC applications benefit from multi core CPU technology and have been able to multiply the computation density by a factor of 2 to 4 and later by 8. This improvement is not enough compared to the computation requirements of today’s applications. This is why people have been looking for new hardware and specialized processors which could give applications gains from 20 up to 100. Specialized processors like GPUs have improved performance at a greater pace than Moore law predicts. They started 10 years ago with a technology using 350nm, 5 million transistors at 75Mhz and now are using 55nm, 700Millions transistors at 800Mhz being able to deliver 512GFlops or more than 3.5GFlops/Watt. This leads to improvements factors of 1.7x/year in transistors count, 1.3x/year in clock speed, 2.0x/year in processing units and 1.3x/year in memory bandwidth. Using such powerful dedicated processors as well as CPU in a highly parallel environment of Multi-core for both is showing the requirement to be able to use in the most efficient way this heterogeneous environment of Hybrid computing. This hardware environment exists and can be used today. The first challenge is on the software development side. Development tools need to integrate heterogeneous programming as well as multi core from the core of their language being able to support code generation on different processors types as well as handling asynchronous behaviors. This comes with compilers and libraries supporting this and being design or extended for it. Obviously those tools need to support multiple hardware platforms to lead to some standards. The second key challenge change is the evolution of buses and bandwidth linking together the different cores of CPU and GPUs. And the way they talk to each other. Fusion projects will address those evolutions in the future by defining new architectures around those processors to improve the data flow between them which will be the key to use all the power available. Cross bar memory controllers will allow GPUs to talk each other very quickly without breaking parallelism. Hyper Transport bus will improve communication between GPUs and CPUs. Finally Multi core GPUs and CPUS on the same die will increase even more the compute density. Different benchmarks and application codes have been used to demonstrate already the benefits so such architecture. We will present SGEMM results as well as different algorithms. The results will highlight the fact that performance is affected by in/out copy of the data on the GPU at the moment and that finer tunings allows huge jump in performance. We will also show that changing the way algorithms have been implemented for CPU to fit GPU architecture adds even more performance gains.

Read the paper · More papers on PaperTik