Tradeoffs in processor-architecture and data-buffer design

J. M. Mulder · 1988

Processor-Memory traffic can be a serious performance constraint for computer systems in general and for closely coupled processors specifically. In a memory-bound system the data traffic becomes a critical performance factor when sufficient instruction buffering is used; when highly encoded instructions are used; or when multiple processors share data memory. Microprocessors are particularly vulnerable to performance degradation because of pin-bandwidth limitations. To relax the memory bandwidth requirements and sustain a high processor throughput, buffering of both instruction and data requests is essential for high performance. This dissertation describes the performance and utility of buffer hierarchies under reference patterns typical for procedural languages. Hardware data-buffering schemes (multiple register sets and stack buffers) can profit significantly from compile-time assistance. Even single register sets perform competitively with these hardware schemes when allocated interprocedurally. In multi-level buffering, first-level (register oriented) and second-level (cache) buffers have distinct purposes. The first-level buffer determines the peak performance of the processor, while the second-level determines the actual performance. Hence, a cache does not hide architectural (first-level buffer) flaws, but emphasizes them, because it reduces the influence of the memory system on performance. The usefulness of caches is a function of the size of the cache and of memory speed. A cache, independent of size, improves the system performance, when memory traffic is the system bottleneck. A cache has to be of significant size to justify implementation when processor performance is a primary design objective and memory is relative fast. A cache connected to an efficient 32-word buffer and 2-cycle memory only reduces the buffer cycles by 25% for 1024 cache words, and by 10% for 128 cache words. The same configuration for a 4-cycle memory yields a cycle reduction of 50% and 30% respectively.

Read the paper · More papers on PaperTik