Hot topic session: How to solve the current memory access and data transfer bottlenecks: at the processor architecture or at the compiler level?

Francky Catthoor, Nikil D. Dutt, Christoforos E. Kozyrakis, foros E. Kozyrakis · 1999

Current processor architectures, both in the programmable and custom case, become more and more dominated by the data access bottlenecks in the cache, system bus and main memory subsystems. In order to provide sufficiently high data throughput in the emerging era of highly parallel processors where many arithmetic resources can work concurrently, novel solutions for the memory access and data transfer will have to be introduced. The crucial question we want to address in this hot topic session is where one can expect these novel solutions to rely on: will they be mainly innovative processor architecture ideas, or novel approaches in the application compiler/synthesis technology, or a mix. 1. Motivation and context At the processor architecture side, previous work has focused on microarchitecture enhancements like intelligent management of cache hierarchies, streaming buffers and value or address prediction, techniques that exploit spatial/temporal locality in memory references. But can this approach provide the memory performance necessary to feed highly parallel processors? We will discuss alternative instruction set architectures that attempt to provide the hardware with explicit information about parallelism in memory, and examine the opportunities and challenges from the emerging combination of multiprocessing and multithreading in a single chip. At the side of the system design technology and compilation for embedded data-dominated multi-media applications, also much evolution is present. We will show that decisions made at this stage heavily influence the final outcome when the appropriate architectural issues of the embedded memories are correctly incorporated. What is more controversial still however is how to deal with dynamic application behaviour, with multithreading and with highly concurrent architecture platforms. 2. Explicitly parallel architectures for memory performance enhancement (Christoforos E. Kozyrakis) While microprocessor architectures have significantly evolved over the last fifteen years, the memory system is still a major performance bottleneck. The increasingly complex structures used for speculation, reordering, and caching have improved performance, but are quickly running out of steam. These techniques are becoming ineffective in terms of resource utilization and scaling potential. In addition, memory system performance will become even more critical in the future, as the processor-memory performance gap is increasing at the rate of 50% per year [28]. Multimedia and embedded applications, that are expected to dominate the processing cycles on future microprocessors, have significantly different memory access characteristics compared to typical engineering workloads [26]. Temporal locality is not always available, leading to poor performance from traditional caching structures. We believe the enhancement of memory system performance requires the synergy of software and hardware techniques. Each system component should be focused on the task it is most efficient with: Compilers and/or run-time tools can view and analyze several hundred lines of source code. They can detect instruction/data parallelism and regular memory access patterns, or transform the code so that parallelism and desired patterns for the given hardware are created (see the two following sections). The processor can utilize the parallelism information to execute concurrently a large number of memory and computation operations using the massive hardware resources available in current and future chips. The runtime (dynamic) information available to the hardware can also be used to apply further optimizations within a small window of instructions. The instruction sets used by CISC and RISC processors today are inherently sequential and hide from the hardware the parallelism and memory access information that is available at various software levels. For example, a group of independent operations are described to the hardware using sequentially ordered instructions. Complex dependency analysis has to be performed in hardware for the parallelism to be ”re-discovered” and utilized. Similarly, a set of sequential or strided accesses to an array will be expressed as a sequence of simple loads or stores to a single memory location. To apply any prefetching or other memory access optimization, the hardware must observe the sequence of addresses and guess the access pattern first. To enable high-performance, yet efficient, processor designs, future instruction set architectures (ISAs) should allow software to pass explicit parallelism and memory access information to the hardware. Instructions should explicitly identify operations that can be issued and executed in parallel. Memory instructions should allow control of memory system features such as cache space allocation, and provide information that describes the access characteristics of the application. This could include access pattern (e.g. sequential or strided), temporal and spatial locality hints, and expected caching behavior at the various levels of memory hierarchy. Using this information, the processor can issue in parallel or overlap a large number of memory accesses, employ aggressive prefetching, and tune the use of the memory hierarchy for maximum performance and efficiency. From the point of view of compilers and run-time tools, explicitly parallel architectures allow fine control of hardware features and create new opportunities for higher-level optimizations. There have been several recent examples of explicitly parallel architectures both from academia and industry. Some examples of different approaches are: The EPIC (IA-64) architecture [27] exposes parallelism to the hardware using VLIW instructions. Memory operations provide explicit hints about the expected caching behavior at each level of the memory hierarchy, which is used for guiding cache space allocation. Software speculation is used to issue memory accesses as early as possible without compromising program correctness. The VIRAM architecture [30] expresses parallelism to hardware in the form of vector operations. Vector memory instructions explicitly specify a large number memory accesses to be issued in parallel, along with their spatial relation (sequential, strided, or indexed accesses). The architecture includes support for software speculation as well. The implementation allows control of address mapping at a per process or a per memory page granularity. The Impulse architecture [25] allows software to describe regular memory access patterns directly to the memory controller. The controller accesses memory in an optimized manner for each pattern, performs prefetching and groups the requested data for maximum caching efficiency. In addition, it provides support for data remapping by the application or the compiler. While instruction set architectures should focus on exposing parallelism to the hardware, processor design should focus on implementations of these ISAs that can tolerate high memory latency, even in the case of poor caching behavior. Increased memory latency is a technology problem unlikely to be solved in the near future. On the other hand, high memory bandwidth is already available through technologies like multi-bank embedded DRAMs, high-performance memory interfaces (Rambus, DDR), and cost-efficient MCM packaging. Processor designs can utilize high memory bandwidth in order to hide the performance penalty of high memory latency. This is in harmony with the characteristics of many multimedia and ecommerce workloads, where throughput is by far more important than latency. Frequently, latency can also be addresseed with software techniques like those presented in the next section. But hardware techniques for Ha hiding latency are still necessary as they can be applied to all applications, regardless of the availability of source code or their suitability to compiler analysis. There are several architectural approaches to utilizing high memory bandwidth. Some of the techniques proposed or used recently include the following: Multithreaded processors, such as Sun MAJC and Compaq Alpha EV-8, attempt to execute concurrently several fine-grain threads. When some thread is blocked due to memory latency (e.g. a cache miss) or lack of parallelism, the hardware switches to issuing instructions from another thread within a couple of clock cycles. Given a large number of fine-grain threads and high memory bandwidth, multithreaded designs can efficiently hide high memory latency. Multiprocessing designs combine two to four processors on the same chip forming a symmetric multiprocessor (IBM Power4, Sun MAJC, Stanford Hydra). Each processor can run a separate execution thread,

Read the paper · More papers on PaperTik