Data Management and Data Processing Support on Array-Based Scientific Data

Yi Wang · OhioLink ETD Center (Ohio Library and Information Network) · 2015

Scientific simulations are now being performed at finer temporal and spatial scales, leading to an explosion of the output data (mostly in array-based formats), and challenges in effectively storing, managing, querying, disseminating, analyzing, and visualizing these datasets.Many paradigms and tools used today for large-scale scientific data management and data processing are often too heavy-weight and have inherent limitations, making it extremely hard to cope with the 'big data' challenges in a variety of scientific domains.Our overall goal is to provide high-performance data management and data processing support on array-based scientific data, targeting data-intensive applications and various scientific array storages.We believe that such high-performance support can significantly reduce the prohibitively expensive costs of data translation, data transfer, data ingestion, data integration, data processing, and data storage involved in many scientific applications, leading to better performance, ease-of-use, and responsiveness.On one hand, we have investigated four data management topics as follows.First, we built a light-weight data management layer over scientific datasets stored in HDF5 format, which is one of the popular array formats.Unlike many popular data transport protocols such as OPeNDAP, which requires costly data translation and data transfer before accessing remote data, our implementation can support server-side flexible subsetting and aggregation, with high parallel efficiency.Second, to avoid the high upfront data ingestion costs of loading large-scale array data into array databases like SciDB, we designed a I would never have been able to complete my dissertation without the guidance of my committee members, help from friends, and support from my family and fiancee over the years.Foremost, I would like to express my sincerest gratitude to

Read the paper · More papers on PaperTik