Management of Very Large Distributed Shared Collections

Marcia J. Bates · Auerbach Publications eBooks · 2011

Scientific data collections are being assembled that contain the digital holdings on which future research is based. The collections are assembled by researchers from multiple institutions, and then accessed by all members of a scientific discipline. The data collections are massive in size, comprising hundreds of terabytes of data (a terabyte is a thousand gigabytes) and tens of millions of files. The software infrastructure that manages these collections must provide not only traditional digital library services, such as indexing, discovery, and presentation, but also preservation services to ensure authenticity and integrity. The types of material in the collections range from digital simulation output generated by scientific applications, to observational data taken by experiments, to real-time sensor data streams from thousands of sensors. Thus the management of scientific data collections requires the integration of capabilities from multiple disparate communities: data grids for sharing data, digital libraries for publishing data, persistent archives for preserving data, and real-time sensor systems for automating the creation of collections.

Read the paper · More papers on PaperTik