Performance enhancement through dynamic scheduling and large execution atomic units in single instruction stream processors
Stephen W. Melvin · 1992
This dissertation demonstrates that through the careful application of hardware and software techniques, general purpose code can be executed more than twice as fast as previously thought possible. Exploiting parallelism is critical to high performance. The type of parallelism focused on in this dissertation is intra-instruction stream, or fine-grained, parallelism. This is the parallelism available within a small dynamic window of instructions executed on a single instruction stream processor. Three mechanisms are analyzed on realistic processors running general purpose code and it is shown that a higher degree of fine-grained parallelism can be exploited than has previously been achieved. It has been suggested that general purpose, or non-scientific, programs have very little parallelism not already exploited by existing processors and can achieve a speedup of at most approximately two. The argument is made that there is little to be gained by complex processor control logic; simple issuing and scheduling mechanisms are sufficient to exploit all the parallelism available. This idea has impeded the development of multiple function unit processors. In this dissertation it is shown that, contrary to this notion, there is a actually a significant amount of unexploited parallelism in typical general purpose programs. Three performance enhancement techniques are analyzed: dynamic scheduling, dynamic branch prediction and basic block enlargement. Dynamic scheduling involves the decoupling of operations within a single instruction word, allowing them to be scheduled independently. Dynamic branch prediction involves the use of speculative execution. Basic block enlargement is a technique that relies on compile time effort and an efficient backup mechanism to exploit parallelism more effectively. It is shown that indeed for narrow instruction words little is to be gained by the application of these three techniques. However, as the number of function units increases, it is possible to achieve speedups of almost five on realistic processors.