Clusters: challenges and opportunities
D.A. Reed · 2004
Summary form only given. The continuum of cluster computing continues to expand, with terascale clusters now in production and petascale clusters in design. How do we manage clusters with tens of thousands of nodes, each with power, communication, processing, and memory constraints? How do we design, package, and support systems with hundreds of thousands of processors in a reliable way? This paper discusses the approaches to computing and communication fault-tolerance and reliability for large-scale clusters. It also sketches some of the technical challenges and opportunities in deploying and supporting large-scale clusters, highlighted by recent developments at NCSA and the U.S. TeraGrid and their application to emerging scientific applications.