Tetrahedrangel: An Energy-Efficient Coalescing Temporal Prefetcher
Sam Ainsworth, Lev Mukhanov · ACM Transactions on Architecture and Code Optimization · 2026
Storing Markov tables inside L3 caches to prefetch temporally correlated addresses can give large speedups in irregular workloads. However, the previous state-of-the-art on-chip temporal prefetcher, Triage, features design inconsistencies and inaccuracies that pose challenges for implementation. We fixed these for our ISCA 2024 design, Triangel, which extends Triage with novel sampling-based methodologies to allow it to be aggressive and timely when the prefetcher is able to handle observed long-term patterns, and to avoid inaccurate prefetches when less able to do so. However, Triangel still suffers from an extremely complex lookup mechanism, with multiple tags compressed inside each L3 cache line, meaning Markov-table entries must be rearranged every time the Markov partition within the L3 changes size. Here we present Tetrahedrangel, which fixes all these issues. It draws in inspiration from newer coalescing temporal prefetchers, such as the OpenXiangShan CMC (X-CMC), which group together multiple prefetch targets under a single lookup address in the Markov table, allowing lookups to be stored simply inside the L3’s own tags. We redesign Triangel’s accuracy-control structures to sit as a small independent module alongside X-CMC’s otherwise very different structures. We also introduce a Remap Table, to virtualize away the aggressive, high-degree prefetching X-CMC must otherwise do to prefetch all 16 targets of each coalesced Markov-table entry. Together with new optimizations to improve accuracy, throttle prefetching, and eliminate redundant Markov-table lookups and trainings, Tetrahedrangel achieves 38% speedup, outperforming Triangel’s 34% and X-CMC’s 25%, while transforming energy-consumption impact: by being highly accurate, frugal in its Markov-table L3 lookups, and more efficient than Triangel in data-storage density, Tetrahedrangel increases combined L3 and DRAM energy by just 7% relative to baseline, compared with 16% for Triangel and 27% for X-CMC, and with far easier paths to deployment by reducing L3 traffic from an impractical 2.2× to just 1.14×.