Distributed Shared Memory
Ratan Kumar Ghosh, Hiranmay Ghosh · 2023
A many-core distributed system consists of multiple multi-core node clusters connected via network on chips (NoCs). Scaling up performance on a many-core system requires careful partitioning and placement of data to extract maximum parallelism. Furthermore, there are technological innovations with specialized hardware which accelerate performance. These include graphics processing unit (GPU), field programmable gateway arrays (FPGAs), and machine/deep learning (ML/DL) accelerators like Google's tensor processing unit (TPU). Simultaneously, intensive research on programming newer architecture led to the discovery of diverse computing abstractions such as MapReduce, Spark, and Dryad for data-parallel applications on these systems. These days, large-scale data-intensive applications like page rank, placement of ads, and social networks are implemented using many-core distributed systems. There is no single abstraction that fits all data-parallel applications. As a result, the applications performing well in one model perform poorly in another. The programmers find the context switching from one abstraction to another difficult, and their productivity suffers. Software distributed shared memory (S-DSM) research saw a resurgence with newer computing platforms. From a programmer's perspective, the shared memory programming model is a natural extension of the uniprocessor memory model on a distributed system. S-DSM implementation is transparent to the programmer. So, it allows the programmers to handle synchronizations in the familiar shared memory model. Their productivity is unhindered by the complexity of message-passing. The biggest challenge in S-DSM implementation is maintaining the consistency of shared variables. Inconsistent values of shared variables lead to a program's behavior being different from the programmer's expectations.