An analysis of multistreamed, superscalar processor architectures

Wayne Yamamoto · 1996

To supply the ever increasing performance demands on computer systems, faster and larger processors are being built. Today's microprocessors exploit higher levels of integration by increasing the cache capacity and adding functional units. As a result, superscalar processors are able to concurrently dispatch an increasing number of instructions on every cycle. However, as the number of functional units increases the probability that a given functional unit is busy decreases. This is due to a lack of instruction level parallelism (ILP) within applications that limit the number of instructions that can be executed in parallel and, thus, the overall performance. While the majority of the effort has been placed on extracting ILP within a single thread of execution, we explore multistreaming, the concurrent execution of instructions from distinct threads, as another technique to increase general purpose, superscalar processor performance. We employ a dynamic instruction interleaving technique where instructions from different threads are dynamically grouped by the processor and simultaneously dispatched to the functional units for execution. The lack of ILP within a thread is overcome by executing instructions from other threads in functional units that a single thread is not able to use. Thus, overall processor performance is significantly increased by using the processor resources more efficiently. In this dissertation, we present an analysis of incorporating multistreaming into a general purpose, superscalar processor. We focus on the effects of functional unit and cache configurations on overall processor performance. Using a combination of analytical models and simulation techniques, we determine the potential and expected performance gains attainable through multistreaming. We measure the impact of thread scheduling on performance and determine scheduling strategies to obtain the highest performance. Finally, we examine implementation costs and show that multistreaming is a cost effective technique for improving processor performance.

Read the paper · More papers on PaperTik