CPACM: A New Embedded Memory Architecture Proposa l
Stephen Richardson · 2000
CPACM, or Combined Paged And Cached Memory, presents a new way to arrange the memory contained in very large caches. CPACM combines the favorable elements of two previous and more well-known schemes, merging the more conventional cachedmemory method with demand paging. This new hybrid scheme consistently performs as well or better than its two predecessors on a wide variety of workloads. This study evaluates CPACM against benchmarks ranging from high end commercial and technical workloads down to very small embedded applications, such as imaging algorithms for printers. Introduction In this paper we present CPACM, or Combined Paged And Cached Memory. In our exploration of applications for chip multiprocessors, we have proposed demand paging for the on-chip level two memory, and compared demand paging with more traditional cache replacement methods [Kelt00]. In that study we found some cases where demand paging performed better than cache, but also some cases where demand paging utterly failed. CPACM combines favorable elements of both methods in an attempt to provide consistently better performance across a variety of workloads. In this study we compare the three memory architectures: demand paging, CPACM and cache. To better understand how results scale for different chip sizes, this study targets two different cost points: one for larger systems and one for smaller, embedded systems. The larger system targets a higher cost point, and its consequently larger die area lets it support more level-two cache memory than the smaller system. The smaller system targets a lower cost point and would be used in smaller workstations and embedded systems. Therefore, the smaller system will be evaluated against a different set of workloads than the larger system. Consider a computer architecture consisting of one or more processors, each of which may have one or more small local memory caches. The processors connect to a large, fast store of local memory backed by a larger, slower remote memory, as shown in Figure 1. Fast Local Memory Slow Remote Memory