Calculation‐Intensive Applications
Michael Di Stefano · 2005
There are many applications that naturally lend themselves to the Grid Architecture; computational intensive applications are one such example. New vocabulary is introduced. Terms such as “parallelizable” and “worklets” are used to describe “Grid-able” applications, i.e. “These applications are parallelizable, they can be broken down into worklets running at the same time independent of each other in a Grid Compute environment”. The common thread connecting this class of application across diverse industries is that they require complex data analysis over increasingly larger data sets. The analytical process traverses, and mines these data sets can require execution times spanning days. Data Grids natively lend themselves to the transient data sets produced by calculation intensive processes such as a Monte Carlo Simulation. These processes generate and leverage vast amounts of interim data that is used throughout the running simulation to produce an end result but are not part of the end result itself. A comparison of the quantity of data input and output from a Monte Carlo Simulation to that of the interim data generated by the running simulation is analogous to an iceberg, with the input and output data representing the tip of the iceberg and the interim data as the majority of the iceberg you do not see. This chapter contrasts the work flows both with and without a Data Grid, discusses two data enhancement techniques, data reuse and data affinity. As with all of the example use cases discussed in this book, the Application Definition Expressions are discussed and the Quality of Service (QoS) vs. Application Requirement Quadrant Graph is generated.