Petascale Computational Systems: Balanced Cyber-Infrastructure in a Data-Centric World

Gordon Bell, Jim Gray, Alexander S. Szalay · Computer · 2006

Computational science is changing to be data intensive. Super-Computers must be balanced systems; not just CPU farms but also petascale IO and networking arrays. Anyone building CyberInfrastructure should allocate resources to support a balanced Tier-1 through Tier-3 design. Computational Science and Data Exploration Computational Science is a new branch of most disciplines. A thousand years ago, science was primarily empirical. Over the last 500 years each discipline has grown a theoretical component. Theoretical models often motivate experiments and generalize our understanding. Today most disciplines have both empirical and theoretical branches. In the last 50 years, most disciplines have grown a third, computational branch (e.g. empirical, theoretical and computational ecology, or physics, or linguistics). Computational Science has meant simulation. It grew out of our inability to find closed form solutions for complex mathematical models. Computers can simulate these complex models. Over the last few years Computational Science has been evolving to include information management. Scientists are faced with mountains of data that stem from four trends: (1) the flood of data from new scientific instruments driven by Moore’s Law – doubling their data output every year or so; (2) the flood of data from simulations; (3) the ability to economically store petabytes of data online; and (4) the Internet and computational Grid that makes all these archives accessible to anyone anywhere exacerbating the replication, creation, and recreation of more data. Acquisition, organization, query, and visualization tasks scale almost linearly with data volumes. By using parallelism, these problems can be solved within fixed times (minutes or hours). In contrast, most statistical analysis and data mining algorithms are nonlinear. Many tasks involve computing statistics among sets of data points in some metric space. Pair-algorithms on N points scale as N. If the data increases a thousand fold, the work and time can grow by a factor of a million. Many clustering algorithms scale even worse. These algorithms are infeasible for terabyte-scale datasets. Computational Problems are Becoming Data-Centric Next generation computational systems and science instruments will generate petascale information stores. The computation systems will often be used to analyze these huge information stores. For example BaBar processes and reprocesses a petabyte of event data today. About 60% of the BaBar hardware budget is for storage and IO bandwidth [1]. The Atlas and CMS systems will have requirements at least 100x higher. The Large Scale Synoptic Telescope (LSST) has requirements in the same range: peta-ops of processing and tens of petabytes of storage. SETI@home and similar projects show that one can do interesting science with an IO-poor environment – but those systems require that the CPU:IO ratio be 100,000 instructions per byte of IO or higher [2]. Cryptography, signal processing, and certain other problem domains have such cpu-intensive profiles, but most other scientific tasks are much more information intensive having CPU:IO ratios well below 10,000:1 in line with Amdahl’s laws.

Read the paper · More papers on PaperTik