Optimizing Access to Scientific Data for Storage, Analysis and Visualization
Latchesar Ionkov · eScholarship (California Digital Library) · 2018
Scientific workflows contain an increasing number of interactingapplications, often with big disparity between the formats of databeing produced and consumed by different applications. This mismatchcan result in performance degradation as data retrieval causesmultiple read operations (often to a remote storage system) in orderto convert the data. In recent years, with the large increase in theamount of data and computational power available there is demand forapplications to support data access in-situ, or close-to simulation toprovide application steering, analytics and visualization.Although some parallel filesystems and middlewarelibraries attempt to identify access patterns and optimize dataretrieval, they frequently fail if the patterns are complex. It isevident that more knowledge of the structure of the datasets at thestorage systems level will provide many opportunities for furtherperformance improvements.For most developers of scientific applications, storing theapplication data, and its particular format on disk, is not anessential part of the application. Although they acknowledge theimportance of the I/O performance, their expertise lies mostly innumerical simulations and the particular models their applicationsimulates. Most of their efforts are spent of ensuring that theit produces correct numerical results. Ideally, they would like to beable to have a library call that reads a subset of the data from storage (nomatter what its format is), and place it in the data structures thesimulation defines in the computer memory. Since the data needs to beanalyzed and visualized, and the data has to be accessible fromthird-party tools, the scientists are forced to know more about thedata formats.In this dissertation we investigate multiple techniques for utilizingdataset description for improving performance and overall dataavailability for HPC applications. We introduce a declarative datadescription language that can be used to define the complete datasetas well as parts of it. These descriptions are used to generatetransformation rules that allow data to be converted between differentphysical layouts on storage and in memory.First, we define the DRepl dataset description language and use it toimplement divergent data views and replicas as POSIX files. Weevaluate the performance for this approach and demonstrate itsadvantages both because of the transparent application use, andcombined performance when the application is combined with analyticsand/or visualization code that reads the data in different format.DRepl decouples the data producers and consumers and the data layoutsthey use from the way the data is stored on the storage system.DRepl has shown up to 2x for cumulative performance when data isaccessed using optimized replicas.Second, we extend the previous approach to the parallel environmentused in HPC. Instead of using POSIX files, the new method allows datato be accessed in larger chunks (fragments) in the way it will be laidout in memory. The developers can define what data structures theyhave in the process' memory and the overall format of the dataset onstorage, and the runtime will automatically take care of transformingthe data between the two. Both the formats in memory and on disk aredescribed with the DRepl language. Replacing the ability for readingthe data as an array of bytes with operations that use descriptions ofthe data structure, provides better opportunities for thestorage system to optimize the access to the persistent data. Theintegration of this technique in Ceph demonstrates the potentialadvantages for this approach. The experiments show performanceimprovements up to 5 times for writes and 10 times for reads, comparedto collective MPI I/O.Third, we explore the future directions of extending the DRepllanguage to support more complex datasets. The additions would allowscientists to use different resolutions for different parts of amulti-dimensional spaces, and define how to transform the data betweenresolutions. The changes would also allow completely abstractdefinitions of datasets not only for continuums, but also forprimitive types like real and integer numbers. The fragments of thedataset that are present in memory or disk will have concretetypes that are compatible with the abstract types used in the dataset.Finally, we provide foundations on how to extend the previousfunctionality to the most complicated data structures used inscientific applications -- unstructured meshes.