Resource management in single-chip multiprocessors

William J. Dally, Kelly A. Shaw · 2005

Technology advances will soon enable billion transistor chips, permitting large quantities of both logic and memory to be placed on a single die. Increasing on-chip wire delays, however, are shrinking the chip area reachable in a single clock cycle. In response, computer architects are redesigning on-chip structures to reduce the distance signals must travel in one clock period. Single-chip multiprocessors are one proposed architecture for dealing with the multi-cycle communication latencies affecting future chips. These chips are organized into a grid of nodes, where each node contains a processor and a portion of the total on-chip memory. Although these nodes function independently, they interact via shared memory through a network with the property that latency increases linearly based on Manhattan distance. This dissertation examines how to efficiently map applications onto single-chip multiprocessors given these chips' constraints and opportunities. The small amount of per-node storage limits how much data can be placed on any given node, while the ample processing resources and high bandwidth, low-latency on-chip communication create a large number of quickly accessible locations where data and threads can reside. In order to achieve high performance on these chips, applications must balance the competing goals of improving locality and of distributing resource demands across the chips' many nodes. I present two symbiotic approaches for managing these chips' resources. The first technique, migration of data and threads, reacts to dynamic resource demands and communication patterns to avoid resource hot-spotting and improve locality. The second approach proactively eliminates communication by executing computation at the location of its most frequently accessed data, its anchor. Finally, I show how these techniques can be used in conjunction with well-established techniques, like caching, to further improve application performance.

Read the paper · More papers on PaperTik