Mobile mpi programs on heterogeneous computational grids
Keshav K. Pingali, Rohit Fernandes · 2006
Utility computing is becoming a popular way of exploiting the potential of computational grids. To take full advantage of utility computing, an application needs to be mobile; that is, it needs to be able to migrate between heterogeneous computing platforms while it is executing. Further, it needs to adapt to the number of available physical processors. At present, there are few high-performance computing applications of this sort, and re-engineering legacy codes to be mobile can take enormous effort. In this dissertation, we propose a system which converts C/MPI codes into mobile programs almost transparently. Our system is based upon three keys ideas: portable application level checkpointing, coordination protocols and over-decomposition. We have developed the Portable Cornell Checkpointing Compiler ( PC3) which automatically transforms a C program to save its own state. On restart, PC3 uses type information to translate program data to an equivalent representation for the new platform. We have designed static and runtime checking mechanisms that can detect when a portable checkpoint may be unsafe for recovery. To checkpoint parallel MPI programs, we have integrated PC 3 with two scalable coordination protocols. These coordination protocols ensure that the collection of checkpoints for the processes composing the application form a distributed snapshot from which the application can correctly recover. The barrier coordination protocol imposes no overheads on the underlying MPI application but is limited to applications that have, globally coordinated phases where there are no in-flight messages. The non-blocking coordination protocol can handle in-flight messages and is applicable to more general MPI programs while keeping overheads low. To provide flexibility of restart on different numbers of processors, we have developed the Virtualized MPI (VMPI) system. VMPI permits MPI applications to be over-decomposed into a larger number of virtual processes than the number of physical processors. The virtual processes are distributed equitably among the physical processors to obtain good load-balance. VMPI leverages fast context switching of user-level threads and efficiently distributes the communication of the application between memory copies and MPI transfers. For most applications, the overheads are low.