Effective Distributed Supercomputing Resource Management for Large Scale Scientific Applications

Seungwoo Rho, Jik‐Soo Kim, Sangwan Kim, Seoyoung C. Kim, Soonwook Hwang · Journal of KIISE · 2015

Nationwide supercomputing infrastructures in Korea consist of geographically distributed supercomputing clusters. We developed High-Throughput Computing as a Service(HTCaaS) based on these distributed national supecomputing clusters to facilitate the ease at which scientists can explore large-scale and complex scientific problems. In this paper, we present our mechanism for dynamically managing computing resources and show its effectiveness through a case study of a real scientific application called drug repositioning. Specifically, we show that the resource utilization, accuracy, reliability, and usability can be improved by applying our resource management mechanism. The mechanism is based on the concepts of waiting time and success rate in order to identify valid computing resources. The results show a reduction in the total job completion time and improvement of the overall system throughput.

Read the paper · More papers on PaperTik