SChISM: Scalable Cache Incoherent Shared Memory
John H. Kelm, Daniel Johnson, Aqeel A. Mahesri, Steven S. Lumetta, Matthew I. Frank, Sanjay Jeram Patel · Illinois Digital Environment for Access to Learning and Scholarship (University of Illinois at Urbana-Champaign) · 2008
This paper motivates and describes a class of accelerator architectures that manage cache coherence in soft-ware to exploit data sharing and communication characteristics present in emerging highly parallel workloads. Based on previous findings and our own studies, we show that our target applications have structure to their com-munication patterns that can be leveraged to move most cache coherence management into software. Replacing hardware cache coherence with software mechanisms allows die area to be reclaimed for more compute resources while also reducing hardware design complexity. Moreover, there is a prevalent programming style for large-scale parallel computation which can be mapped into a low-level task-based programming model that manages coherence in software without sacrificing usability and performance. We also observe that the main benefit of hardware cache coherent systems is not for supporting data-parallel applications, but rather for implementing preemptive multitasking operating systems where migratory data and global locking constructs must be supported efficiently. To support the system software features that are re-quired for applications running on an accelerator platform, we demonstrate an implementation of the Rigel Task Model. The Rigel Task Model is a low-level programming model that performs work distribution, task scheduling, software-enforced cache coherence, and synchronization in software, with limited specialized hardware, for an accelerator architecture without hardware cache coherence. 1