A Method of Porting Multi-Threaded Programs to Cluster

Zha Li · Chinese Journal of Computers · 2002

Scientific computing is moving from mainframes to distributed clusters, and extends to Internet and Grid. The main focuses of cluster computing are parallel algorithm, communication and scheduling in recent years. The user interface and programming style of the application running on cluster systems is ignored or not seriously considered. Our effort is on this point accordingly. As a presumption, the multi threaded program which will be ported to cluster must run correctly on single node. Our approach is to modify the binary code of that program by ELF Rewriter directly, and distribute the computation loads among the whole cluster system without touching the source code. The ELF Rewriter injects a stub code segment into the host program, responsible for the communication between host program and Task Dispatcher, and host program and Communication Server. When original shared library calls in the host program are captured, ELF Rewriter just redirects them to the newly injected stub code segment. We design the Communication Client and the Task Executing Agent on each node to transfer and execute computation loads. When running on cluster, there is one Master node for the host program, and the others are Worker nodes. The Master node no longer executes the CPU intensive library calls in fact, but only communicates with Worker nodes. We develop a testing framework of this computing model, and propose and discuss a Master Worker (Task Farming) computation and communication model with corresponding scheduling policy. Finally, a multi threaded matrix multiply program based on BLAS v3 library is implemented, and being ported to cluster environment so as to verify the feasibility and efficiency of the model. By the ELF executable PLT redirection, we do not need changing the source code to manually distribute the loads. It is automatically distributed to all nodes. The concurrency and synchronization of the host multi threaded program are utilized, no extra synchronization is needed. We just by this mean to provide a user transparent cluster range parallel mechanism.

Read the paper · More papers on PaperTik