Fetching instruction streams

Alex Ramírez, Oliverio J. Santana, Josep-L. Larriba-Pey, Mateo Valero · 2002

Fetch performance is a very important factor because it effectively limits the overall processor performance. How-ever, there is little performance advantage in increasing front-end performance beyond what the back-end can con-sume. For each processor design, the target is to build the best possible fetch engine for the required performance level A fetch engine will be better if it provides better per-formance, but also if it takes fewer resources, requires less chip area, or consumes less power. In this paper we propose a novel fetch architecture based on the execution of long streams of sequential instructions, taking maximum advantage of code layout optimizations. We describe our architecture in detail, and show that it re-quires less complexity and resources than other high perfor-mance fetch architectures like the trace cache, while provid-ing a high fetch performance suitable for wide-issue super-scalar processors. Our results show that using our fetch architecture and code layout optimizations obtains 10 % higher performance than the EV8 fetch architecture, and 4 % higher than the FTB architecture using state-of-the-art branch predictors, while being only 1.5 % slower than the trace cache. Even in the absence of code layout optimizations, fetching instruc-tion streams is still lO % faster than the EV8, and only 4% slower than the trace cache. Fetching instruction streams effectively exploits the spe-cial characteristics of layout optimized codes to provide a high fetch performance, close to that of a trace cache, but has a much lower cost and complexity, similar to that of a basic block architecture. 1.

Read the paper · More papers on PaperTik