Simulation Runner: A Cloud-Based Parallel and Distributed HPC Platform
Zhenbang Liu, Hengming Zou, Wenming Ye · 2015
Scientific workloads often require scalable computing resources. The traditional HPC paradigm provides large scale computing resources but requires IT specialists to buy and maintain large computer infrastructure in their organization. With the emergence of cloud computing, more and more scientists and engineers are moving their computing workloads from local clusters to public cloud. However, such a migration often requires experience both in understanding cloud architecture and cloud development skills, which can impede the efficiency and success of migration. This paper presents a cloud-based parallel and distributed computing platform to support scientific computing. With the Simulation Runner, scientists only need to submit their embarrassingly parallel scientific applications with appropriate arguments or parameters. There is no need to rebuild migrate cloud applications or purchase large computer infrastructures by themselves. The system also uses a file caching mechanism to improve performance and reduce the costs of data-intensive computing tasks. A 44X speedup was resulted from using 128 computing instances compared with premise high performance desktop computers. A comparison study on the performance of file cache is given at the end of the paper.