Process migration in heterogeneous distributed computing systems
Mallik V. Yalamanchili · 1997
Providing for valuable features like heterogeneous processing, load balancing, and crash tolerance in heterogeneous distributed systems necessitates heterogeneous process migration. This migration involves transferring some significant subset of the execution state of a process from a source machine to a heterogeneous destination machine and translating it into a valid state on the destination machine so that execution can resume on the destination machine from the point where it was stopped on the source machine. This research deals with one of the main problems of heterogeneous process migration--architectural heterogeneity, specifically heterogeneity in register cardinality, which is defined as difference in the number of registers on two machines. This heterogeneity is directly reflected in the process states of the two machines and in the machine code optimized for the machines, which decomposes the problem further into two sub-problems: translating the process state and finding corresponding migration points in object code. These two problems are simultaneously addressed in this research by using a machine-independent optimized program as a platform for migration. At compilation time, when machine-dependent optimized programs are generated using a machine-independent optimized program for both of them, information on their corresponding migration points, the active sets associated with the migration points, and the correspondences among the members of those active sets are generated. These active set members are variables or registers used in the object code that provide the necessary and sufficient information for computation to progress beyond a point in a program. At runtime, through lookups, this information is used to accomplish migration-point and execution-state translation between the source machine's object code and the destination machine's object code. The principal advantage of this approach is that it doesn't require any recompilation during migration and execution on the destination machine can resume at exactly the same point where it was stopped on the source machine. This schema can be incorporated easily into existing compilers and operating systems, and state-translation can be accomplished quickly and accurately without disturbing the load on the distributed system. It can serve as framework for providing solutions for other types of architectural hindrances to heterogeneous process migration.