Dealing in Practice with Memory Hierarchy Effects and Instruction Level Parallelism

Sid Ahmed Ali Touati, Benoît Dupont de Dinechin · 2014

The chapter provides the memory disambiguation mechanisms in some high-performance processors. Such mechanisms, coupled with load/store queues in out-of-order processors, are crucial to improving the exploitation of instruction-level parallelism (ILP), especially for memory-bound scientific codes. The chapter discusses the cache effects optimization at instruction level for embedded very long instruction word (VLIW) processors. The introduction of caches inside processors provides micro-architectural ways to reduce the memory gap by tolerating long memory access delays. The chapter presents a backend code optimization for tolerating non-blocking cache effects at the instruction level. It shows how to study the dynamic behavior of memory request processing, and provides examples on three superscalar processors. The chapter describes the aspect of memory hierarchy, which is cache misses penalties. It presents a method to reduce processor stalls due to cache misses in presence of non-blocking cache architectures.

Read the paper · More papers on PaperTik