Non-Preconditioned Conjugate Gradient on Cell and FPGA Based Hybrid Supercomputer Nodes
David H. DuBois, Andrew DuBois, Thomas M Boorman, Carolyn Connor · 2009
This work presents a detailed implementation of a double precision, non-preconditioned, conjugate gradient algorithm on a Roadrunner heterogeneous supercomputer node. These nodes utilize the Cell Broadband Engine Architecturetrade in conjunction with x86 Opterontrade processors from AMD. We implement a common conjugate gradient algorithm, on a variety of systems, to compare and contrast performance. Implementation results are presented for the Roadrunner hybrid supercomputer, SRC Computers, Inc. MAPStation SRC-6 FPGA enhanced hybrid supercomputer, and AMD Opteron only. In all hybrid implementations wall clock time is measured, including all transfer overhead and compute timings.