Macro-op scheduling and execution
Ilhyun Kim, Mikko H. Lipasti · 2004
High-performance microprocessors must attain two elusive goals that are often at odds with each other: high operating frequency that demands a minimum amount of logic per pipeline stage, and a high degree of concurrency in the form of instruction-level and memory-level parallelism, which tends to increase the amount of activity required to exe-cute each instruction. The logic required to implement all of this complexity can be broken down into more and more pipeline stages to achieve higher operating frequency, but this comes at the cost of additional power, pipeline latch overhead, and erosion in IPC due to increased branch and scheduling penalties. One promising approach to overcoming this limitation and simplifying the control logic overhead in an out-of-order processor is to move from the conventional instruction-level processing towards coarse-grained instruction processing that amortizes these over-heads over a set of two or more instructions. This thesis proposes and evaluates two microarchitectural techniques that perform coarse-grained instruction processing in super-scalar, out-of-order processors: Macro-op Scheduling and Macro-op Execution.