Developing metadata standards for scientific data reuse in NCSA's distributed grid architecture
Joe Futrelle · 2002
The unprecedented availability of network bandwidth and storage has brought about an explosion of on-line scientific data collections. In virtually all scientific fields, data is growing rapidly in volume and complexity. Managing distributed data collections is rapidly becoming one of science's key challenges, particularly for emerging interdisciplinary fields where it is critical for a given study to have access to multiple, heterogeneous collections. Scientific data models are especially difficult to integrate because of their wide variety, high granularity, large storage requirements, and open-ended use requirements. NCSA is fortunate to have an opportunity, through its applications technologies teams and its close relationships with government organizations such as NASA and the National Cancer Institute, to investigate new tools and strategies for integrating distributed scientific data collections. Through the development of application-specific use scenarios, the authors have attempted to sort their the common data modeling issues for a range of applications, including information retrieval, analysis, and visualization. In particular, they have been interested in metadata since its primary use is to enable data access and interoperability.