Temporal memory streaming

Thomas F. Wenisch · 2007

While device scaling has led to continued processor performance improvement, scaling trends in DRAM technology have favored improving density over access latency. As a result, processors in modern servers spend much of execution time stalled on long-latency memory accesses. The conventional approach to latency tolerance—enlarging the on-chip cache hierarchy as transistor budgets scale—is providing diminishing returns because today's multi-megabyte caches already capture available locality. Commercial server applications present a particular challenge for memory system design because current prefetching/streaming approaches are often ineffective on the irregular data structures and dependent miss chains characteristic of these applications. To further improve server performance, architects must design mechanisms that issue memory requests earlier and with greater parallelism in the face of complex access patterns. Despite their complexity, commercial applications nonetheless execute repetitive code sequences, which give rise to recurring data structure traversals. As a result, memory addresses are temporally-correlated—addresses accessed near one another in time often recur together. By recording temporally-correlated cache miss addresses and using the recorded information to predict

Read the paper · More papers on PaperTik