IBM Research Report The Cell Broadband Engine: Exploiting Multiple Levels of Parallelism in a Chip Multiprocessor

Michael K. Gschwind · 2006

AbstractAs CMOS feature sizes continue to shrink and traditional microarchitectural methods for delivering high per-formance (e.g., deep pipelining) become too expensive and power-hungry, chip multiprocessors (CMPs) become anexciting new direction by which system designers can deliver increased performance. Exploiting parallelism in suchdesigns is the key to high performance, and we find that parallelism must be exploited at multiple levels of the sys-tem: the thread-level parallelism that has become popular in many designs fails to exploit all the levels of availableparallelism in many workloads for CMP systems.We describe the Cell Broadband Engine and the multiple levels at which its architecture exploits parallelism:data-level, instruction-level, thread-level, memory-level, and compute-transfer parallelism. By taking advantage ofopportunities at alllevelsof the system, this CMPrevolutionizes parallel architectures todeliver previously unattainedlevels of single chip performance.We describe how the heterogeneous cores allow to achieve this performance by parallelizing and offloadingcomputation intensive application code onto the Synergistic Processor Element (SPE) cores using a heterogeneousthread model with SPEs. We also give an example of scheduling code to be memory latency tolerant using softwarepipelining techniques in the SPE.

Read the paper · More papers on PaperTik