High Performance Communication Framework for Large Scale Workflows
X. Wang, Uwe Küster, MICHAEL M. RESCH, Erich Focht · Civil-comp proceedings · 2011
In distributed workflow systems, high volumes of intermediate data, e.g. hundreds of terabytes, have to be transferred between interconnected tasks. In many cases, due to the lack of a common transferring interface and the know-how to perform highperformance communications, scientific researchers and engineers have to resort in their workflows to writing result data of prior tasks to a local disk, shifting them via network and later fetching the data to the post-processing tasks. Workflows with such transferring methods encounter bottleneck problems, since I/O systems slow down significantly and even may crash by handling large amounts of requests. In addition, the general tendency shows that the amount of data that has to be transferred among complex workflow tasks may increase essentially. The storage limitations and the expensive I/O operators lead to develop an efficient and scalable technique to help researchers to execute scientific workflows on HPC systems. In this paper we present a common framework that provides highly efficient and scalable communications between distributed inter-dependent tasks within workflows by using the interconnection network. Our system is a simple and reliable tool to relieve the researchers from network programming and data management. Another major advantage of our approach is that, it supports various applications, from in-house codes to commercial softwares, without modification to the existing applications. Moreover, it provides direct and fully automatic data streaming among distributed tasks. Experimental results on a biomechanical workflow confirm that the framework has the potential to greatly improve performance and scalability.