The use of copy-on-reference in a process migration system
Edward R. Zayas · 1987
Process migration is a valuable tool in a distributed programming environment. Two factors have conspired to discourage efficient implementations of this facility. First, it has been difficult to design systems that offer the necessary name and location transparency at a reasonable cost. Also, it is often prohibitively expensive to copy the large virtual address spaces found in modern processes to a new machine, given the narrow communication channels available in such systems. This dissertation examines the use of lazy address space transfers when processes are migrated to new sites. An IOU for all or part of the process memory is transferred to the remote location. Individual memory pages are copied over the network in response to attempts by the transplanted process to touch areas for which it holds an IOU. The S sc PICE environment developed at Carnegie-Mellon has been augmented to provide a process migration facility that takes advantage of such a copy-on-reference scheme. The underlying Accent kernel's location-independent IPC mechanism is integrated with its virtual memory management to supply the necessary transparency and the ability to transmit data in a lazy fashion. Study of the testbed system reveals that copy-on-reference address space transmission improves migration effectiveness (performance). Relocations occur up to a thousand times faster, with transfer times independent of process size. Since processes access a small portion of their memory in their lifetimes, the number of bytes transferred between machines drops by up to 96%. Message-handling costs are lowered by up to 94%, and are more evenly distributed across the remote execution. Without special tuning, faulting in a remote page took only 2.8 times longer on average than accessing a page on the local disk. Page prefetch and explicit transfer of resident pages are shown to be helpful in certain situations. Implementation, instrumentation and study of the testbed provides more useful information than is possible through a simulation. The results are applicable to a family of related systems, and point to improvements in distributed operating system design.