REMOTE UNIX TURNING IDLE WORKSTATIONS INTO CYCLE SERVERS
Michael Litzkow · 1992
A computing environment consisting of workstations connected by a local area network is now common. Often these workstations are assigned to individual users, and thus represent a significant unused resource when those individuals are not working. Remote Unix, (RU) uses these idle workstations to execute compute-bound jobs in the background. Users submit jobs to RU from their own workstations. The jobs are queued, and eventually executed remotely on idle workstations. When their jobs have completed, users are notified by mail. The owner of a workstation has absolute priority over RU jobs. When the owner initiates interactive or other non-RU work, any RU job is automatically checkpointed and restarted on another workstation. RU jobs may go through any number of checkpoints before eventual completion. Finding idle workstations, handling of remote system calls, and checkpointing when necessary are all handled by the RU software without intervention from users. In addition to providing ‘‘free’’ cycles, RU allows completion of very long running jobs which might otherwise be aborted by system crashes and shutdowns. The longest running job so far has completed successfully after accumulating 60 CPU days over a 3 month period. This paper describes the computing environment for which RU was designed, and current limitations on the class of jobs which it can handle. The three main components of RU — remote system call handling, a general UNIX† checkpointing facility, and distributed spooling and control — are each discussed. The current version of RU supports only single process jobs. Possible extensions are discussed.