Exploiting data similarity in multicore systems
Frederic T. Chong, Susmit Biswas · 2010
While microprocessor designers turn to multicore architectures to sustain performance expectations, the dramatic increase in parallelism of such architectures will put substantial demands on off-chip bandwidth and make the memory wall more significant than ever. We argue that one profitable application of multicore processors is the execution of many similar instantiations of the same program. We identify that this model of execution is used in several practical scenarios and term it as “multi-execution.” We also find that each such instance often utilizes very similar data. We observe this computing model more prevalent in high performance computing (HPC) domain where many MPI (message passing interface) tasks solve fragments of a large problem and infrequently communicate with each other. In the HPC domain, the limiting factor is the amount of physical memory as the compute nodes often lack disk storage. In a conventional memory system, a process manages its own data independently. I propose that significant improvement in performance or the capability to solve larger problems is achievable by leveraging the similarity in data across processes in multicore nodes through hardware or software means. I propose a hardware assisted technique and a purely software based technique. The hardware approach introduces “Mergeable cache” which maintains a single copy of identical data blocks in cache, and thereby, increases cache capacity, thereby leading to a speedup by a factor of 2.1 on average for the multi-execution workloads. In the software approach, the target being a transparent user level solution, I developed a memory allocation library, “SBLLmalloc”, that uses shared memory to reduce duplicate pages from MPI tasks transparently and reduces the memory footprint of MPI tasks in every node to run 20% larger problems than are solvable using the same resources otherwise.