Data-intensive computing and digital libraries
Reagan Moore, Thomas A. Prince, Mark H. Ellisman · Communications of the ACM · 1998
How to automate management of the flood of scientific data being collected in astronomical and neuroscience projects.Computational science is expanding to include not only analyses based on simulations but those requiring manipulation of large data collections.The need for data-intensive computing is being driven by the massive amounts of data now available in various scientific disciplines.As a consequence, supercomputing systems have to incorporate data and information-handling technologies to manage these very large data sets.And as the amount of data in storage environments, such as digital libraries, increases, it will be necessary to use supercomputers to analyze their holdings.The coevolution of supercomputer and digital library technologies will therefore form the basis for future scientific applications and information-analysis activities.For supercomputers, the creation of information and its organization and use in future computations is the ultimate goal.Infrastructure development by the Data Intensive Computing Environments group in NPACI represents a first step in this direction.