The Armada framework for parallel I/O on computational grids
Ron A. Oldfield · 2002
An exciting trend in high-performance distributed computing is the development of widely-distributed networks of heterogeneous systems and devices, known as computational grids. Grid applications use high-speed networks to logically assemble collections of resources such as scientific instruments, supercomputers, databases, and so forth. One important challenge facing grid computing is efficient parallel I/O for data-intensive grid applications. Data-intensive grid applications are particularly challenging because they require access to large (terabyte-petabyte) remote data sets and often have computational requirements that can only be met by high-performance supercomputers. In addition, data is often stored in “raw” formats and requires significant preprocessing or filtering before the computation can take place. Such applications exist in seismic processing, climate modeling, physics, astronomy, biology, chemistry, and visualization. In this report, we present the Armada framework [OK01] for building I/O-access paths for data-intensive grid applications. We designed Armada to allow grid applications to efficiently access data sets distributed across a computational grid, and in particular to allow the application programmer and the dataset provider to design and deploy a flexible network of application-specific and dataset-specific functionality across the grid. Using the Armada framework, grid applications access remote data sets by sending data requests through a graph of distributed application objects. The graph is called an “armada” and the objects are called “ships”. Figure 1 shows a simple armada for an application accessing applying a preprocessing operator to a distributed data set. We expect most applications to access data through existing armadas constructed by a data set provider; however, it is also possible for the application to extend existing armadas with applicationspecific functionality or to construct entire armadas from scratch. The armada encodes the programmer’s interface, data layout, caching and prefetching policies, interfaces to heterogeneous data servers, and most other functionality provided by an I/O system. The application sees an armada as an object providing access to a specific type of data through a high-level interface. One use of Armada, for example, is to construct complicated data sets on top of legacy files and databases. API storage