Analysis of cache performance in vector processors and multiprocessors
J.D. Gee · 1993
This dissertation examines the memory reference behavior and cache performance in vector processors and multiprocessors. In both areas, cache memories are either extremely important or will become important as processor speeds and system size continue to increase. Vector machines have avoided caches in the past, but there is renewed interest as memory latencies continue to increase relative to processor speeds, and as it becomes feasible to implement very large caches. Multiprocessors have typically depended on caches to reduce demand on main memory, but then require the use of cache consistency (coherency) protocols to manage the caching of shared writable data. The development of high performance protocols continues to be a topic of great interest. Vector processor caches are evaluated by first examining the reference locality and cache performance in a number of vector applications. Vector applications are found to contain significant locality, and measured cache miss ratios are shown to be extremely low for large caches of several megabytes or more. Accurate timing simulators of vector machines equipped with large data caches were then developed and used to measure performance across a wide range of applications. Results suggest that large data caches can double the performance of a vector machine. Multiprocessor caches are studied by examining and characterizing sharing patterns in a large and varied sample of parallel address traces. Results from this analysis are used to identify features of existing cache-consistency protocols that are most useful in improving performance, and also to aid the development of improved cache consistency protocols. Efficient algorithms were then used to simulate a wide variety of consistency protocols over a wide range of implementation parameters. I found that program characteristics often determine which protocols perform best, although protocols which can adapt to different sharing patterns perform well in most circumstances. A more important result is that shared memory multiprocessors do not scale well on existing programs, as overhead to maintain consistent caches is the dominant component of the total execution time. We believe that parallel programs will have to be rewritten to make effective use of shared memory multiprocessors.