A Near-Memory Processor for Vector, Streaming and Bit Manipulation Workloads

Ming‐Liang Wei, Marc Snir, Josep Torrellas, R. B. Tremaine, Thomas M. Siebel, N. Goodwin · Illinois Digital Environment for Access to Learning and Scholarship (University of Illinois at Urbana-Champaign) · 2005

Many important applications exhibit poor temporal and spatial locality and perform poorly on current commodity processors, due to high cache miss rates. In addition, they sometimes need to perform expensive bit manipulation operations that are not efficiently supported by commodity instruction sets. To address this problem, this paper proposes the use of a heterogeneous architecture that couples on one chip a commodity microprocessor together with a coprocessor that is designed to run well applications that have poor locality or that require bit manipulations. The coprocessor supports vector, streaming, and bit-manipulation computation. The coprocessor is a blocked-multithreaded narrow in-order core. It has no caches but has exposed, explicitly addressed fast storage. A common set of primitives supports the use of this storage both for stream buffers and for vector registers. We simulated this coprocessor using a set of 10 benchmarks and kernels that are representative of the applications we expect it to be used for. These codes run much faster, with speedups of up to 18 over a commodity microprocessor, and with a geometric mean of 5.8. 1.

Read the paper · More papers on PaperTik