Application programming on a shared memory multicomputer
Todd Poynor, Tom Wylegala · 2000
This paper describes our experience to date in the HP Labs MultiComputer Systems (MCS) project investigating issues involved in writing applications for a shared memory multicomputer, with an emphasis on fault containment for various types of hardware and software failures. Shared memory multicomputers hold considerable promise for the commercial marketplace as modular architectures that transcend SMP scaling bottlenecks while preserving SMP-like memory load/store programming models. Many multicomputer platforms today do not fully deliver these benefits because most resources are partitioned and shared memory is limited to inter-node communication via message passing, as in a shared-nothing cluster. Those multicomputer platforms that we are aware of that do allow access to global resources have limited support for containing faults from propagating across nodes, leaving multicomputers at a disadvantage when compared to conventional clusters in this regard