Scheduling distributed data-intensive applications on global grids

Srikumar Venugopal · Minerva Access (University of Melbourne) · 2006

The next generation of scientific experiments and studies are being carried out by large collaborations of researchers distributed around the world engaged in analysis of huge collections of data generated by scientific instruments.Grid computing has emerged as an enabler for such collaborations as it aids communities in sharing resources to achieve common objectives.Data Grids provide services for accessing, replicating and managing data collections in these collaborations.Applications used in such Grids are distributed data-intensive, that is, they access and process distributed datasets to generate results.These applications need to transparently and efficiently access distributed data and computational resources.This thesis investigates properties of data-intensive computing environments and presents a software framework and algorithms for mapping distributed data-oriented applications to Grid resources.The thesis discusses the key concepts behind Data Grids and compares them with other data sharing and distribution mechanisms such as content delivery networks, peer-to-peer networks and distributed databases.This thesis provides comprehensive taxonomies that cover various aspects of Data Grid architecture, data transportation, data replication and resource allocation and scheduling.The taxonomies are mapped to various Data Grid systems not only to validate the taxonomy but also to better understand their goals and methodology.The thesis concentrates on one of the areas delineated in the taxonomy -scheduling distributed data-intensive applications on Grid resources.To this end, it presents the design and implementation of a Grid resource broker that mediates access to distributed computational and data resources running diverse middleware.The broker is able to discover remote data repositories, interface with various middleware services and select suitable resources in order to meet the application requirements.The use of the broker is illustrated by a case study of scheduling a data-intensive high energy physics analysis application on an Australia-wide Grid.The broker provides the framework to realise scheduling strategies with differing objectives.One of the key aspects of any scheduling strategy is the mapping of jobs to the appropriate resources to meet the objectives.This thesis presents heuristics for mapping jobs with data dependencies in an environment with heterogeneous Grid resources and multiple data replicas.These heuristics are then compared with performance evaluation metrics obtained through extensive simulations.This is to certify that (i) the thesis comprises only my original work, (ii) due acknowledgement has been made in the text to all other material used, (iii) the thesis is less than 100,000 words in length, exclusive of table, maps, bibliographies, appendices and footnotes.

Read the paper · More papers on PaperTik