Prefetch-directed Scheme for Accelerating Memory Accesses of PCIe-based I/O Subsystem

Wuzhe Wang, Peiyong Zhang · 2023

In modern high-performance CPUs, the Last Level Cache (LLC) is an important resource, commonly shared by various cores. Technologies like Intel’s Data Direct I/O (DDIO) are also proposed to allocate LLC resources for high-speed I/O, thereby enhancing the I/O performance. We argue that data prefetching can more aggressively improve the LLC utilization for the I/O subsystem. In this paper, we introduce a prefetching scheme for the PCIe-based I/O subsystem integrated a multi-core CPU design. Both the prefetcher and PCIe controller are connected to a coherent interconnect and share the LLC with other cores. The prefetcher monitors bus transactions of the PCIe controller and records its memory accesses split by different pages in the form of delta sequences. When a new access sequence matches the recorded one, prefetch candidates are generated in a recursive manner, called lookahead mechanism, and eventually transformed into cache stashing requests sent to the interconnect to prefetch data from off-chip DDR memory to LLC, hiding the long latency of memory accesses. Based on register transfer level (RTL) design, cycle-accurate full-system simulation results show that, for some typical memory access patterns of PCIe devices, our prefetch-directed scheme can on average reduce latency by 11.3% and increase bandwidth by 12.6%. Logic synthesis results based on an advanced process node also indicate that the area overhead of the prefetcher is only 6.7% of the integrated PCIe subsystem, which is well suited for the chip design of large-scale CPUs.

Read the paper · More papers on PaperTik