The second trans-Pacific Grid Datafarm testbed and experiments for SC2003
Osamu Tatebe, Haruo Ogawa, Y. Kodoma, Tomohiro Kudoh, Satoshi Sekiguchi, Satoshi Matsuoka, Kento Aida, Taisuke Boku, Mitsuhisa Sato, Y. Morita, Yoshihiro Kitatsuji, John Michael Williams, J. B. Hicks · 2004
The Grid Datafarm architecture is designed for global petascale data-intensive computing. It provides a global parallel file system (Gfarm file system) with online petascale storage, scalable I/O bandwidth, and scalable parallel processing by federating thousands of local file systems in a grid of clusters securely using Grid security infrastructure. One of features is that it manages file replicas in filesystem metadata for fault tolerance and load balancing. Here, we present an overview of our planned experiment performed as the SC2003 Bandwidth Challenge at the Supercomputing 2003 site in Phoenix, Arizona, USA. In the experiment, five clusters in Japan and three clusters in US comprise a Gfarm file system, on which world-wide largescale data analysis is performed. In the Gfarm file system, a file is dispersed in several cluster nodes, each of which is replicated independently and in parallel by multiple third-party transfers between multiple cluster nodes. For the Challenge, terabyte-scale experimental data is replicated between US and Japan via APAN/TransPAC and SuperSINET (about 10,000 km or 6,000 miles). At the workshop we present the full detail of the experiment.