High performance distributed data reduction and analysis with the netCDF Operators (NCO)
CS Zender, DL Wang · eScholarship (California Digital Library) · 2007
3B.4 HIGH PERFORMANCE DISTRIBUTED DATA REDUCTION AND ANALYSIS WITH THE NETCDF OPERATORS (NCO) Charles S. Zender ∗ and Daniel L. Wang University of California, Irvine 1. INTRODUCTION Gridded geoscience model and sensor datasets present an interesting set of challenges for researchers and the data portals that serve them (Foster et al., 2002). Many geoscience disciplines have transitioned or are transition- ing from data-poor and simulation-poor to data-rich and simulation-rich (NRC, 2001). A software ecosystem has evolved to help researchers exploit this transition with fast data discovery, aggregation, analysis, and dissemi- nation techniques (e.g., Domenico et al., 2002; Cornillon et al., 2003). In this ecosystem are the netCDF Operators (NCO)—software for manipulation and analysis of grid- ded geoscience data stored in the self-describing netCDF format. NCO is used in several niches in geoscience data analysis workflow (Woolf et al., 2003), because its func- tionality is independent of and complementary to data dis- covery, aggregation, and dissemination. The netCDF Operators have evolved over the past decade to serve research the needs of individual re- searchers and data-centers for fast, flexible tools to help manage netCDF-format datasets. The NCO User’s Guide (Zender , 2006a) fully documents NCO’s functionality and calling conventions. Zender (2006b) describes NCO’s design philosophy, primary features, relation to other geo- science data analysis software, and future plans. Zen- der and Mangalam (2007) describe the core NCO arith- metic algorithms and their theoretical and measured scal- ing with dataset size and structure. Current research pro- vides NCO with advanced parallel computing techniques at two distinct levels. At the low level, all NCO arith- metic operators are parallelized to throughput on shared memory and distributed memory clusters (Zender et al. 2007, manuscript in preparation). High level analysis scripts of multiple NCO commands benefit from out new dependency-analysis engine which automatically detects and parallelizes basic blocks, returning intermediate files only as needed (Wang et al., 2006, manuscript in prepa- ration). This extended abstract summarizes novel fea- tures in NCO’s design, arithmetic algorithms, and low- and high-level parallelization. 2. DESIGN Traditional scientific data processing works with an intra- file paradigm where users open one or a few files to read and manipulate one or a few variables at a time. The intra-file paradigm works well in cases where all the perti- Corresponding author address: Charles S. Zender, Dept. of Earth System Science, University of California, Irvine, Irvine, CA 92697-3100; e-mail: [email protected]. nent data reside in a few files, and the processing of each variable is unique and requires hand-coding. In large geoscience applications data storage requirements may dictate that relevant data be spread over multiple files. Level one satellite data, for example, are often stored in a file-per-day or file-per-orbit format. Data produced by geophysical time-stepping models is usually output ev- ery time-step or as a series of time-averages. Climate models usually archive data once per simulated day or month in multi-year or multi-century simulations. NCO supports an inter-file paradigm for situations where the intra-file paradigm is unwieldy. NCO abides by five guidelines that have proven their value when processing large numbers of geophysical datasets: 1. Files behave as an elemental data unit. Unless specifically requested otherwise, NCO applies the same operation to all variables (or attributes) in a file. Manipulating (e.g., adding, subtracting) entire geophysical states as represented by the collection of variables in a file is as easy as manipulating a sin- gle variable in a traditional data analysis language. When the “process all variables” paradigm is com- bined with UNIX command-line globbing of multi- ple files, NCO effectively subsumes two problematic loops (loops over files and over variables) of large scale data-processing into one command. 2. Files processed sequentially are usually homoge- neous. NCO assumes the structure of each file (i.e., the fields present and their dimensions) are identical to the structure of the first file in the sequence. NCO allows the record dimension (usually time) length and number of variables to change between files, but not the ranks of variables. 3. An audit trail that tracks data provenance and pro- cessing history is desirable for both the data ana- lyst and their colleagues who receive the processed data. For analysis involving multi-file sequences, the metadata in the first file, along with a list of the other files, adequately preserves the processing history. By convention, NCO keeps this information in the history attribute (Rew et al., 2005). 4. There is value in maintaining the distinctions and associations between dimensions, coordinates, and variables during data analysis. Unless otherwise specified, NCO automatically attaches coordinate data (i.e., dimension values) to variables it transfers. 5. Tools should treat data as generically as possible, and impose no software limitations on data dimen- sionality, size, type, or ordering.