Data Consolidation and Information Aggregation in Grid Networks
Panagiotis C. Kokkinos, Emmanouel Varvarigos · InTech eBooks · 2011
IntroductionGrids consist of geographically distributed and heterogeneous computational and storage resources that may belong to different administrative domains, but are shared among users by establishing a global resource management architecture.A variety of applications can benefit from Grid computing; among them data-intensive applications that perform computations on large sized datasets, stored at geographically distributed resources.In this context, we identify two important issues: i) data consolidation that relates to the handling of these data-intensive applications and ii) information aggregation, which relates to the summarization of resource information and the provision of information confidentiality among the different administrative domains.Data consolidation (DC) applies to data-intensive applications that need more than one pieces of data to be transferred to an appropriate site, before the application can start its execution at that site.It is true, though, that an application/task may not need all the datasets at the time it starts executing, but, it is usually beneficial both for the network and for the application to perform the datasets transfers concurrently and before the task's execution.The DC problem consists of three interrelated sub-problems: (i) the selection of the replica of each dataset (i.e., the data repository site from which to obtain the dataset) that will be used by the task, (ii) the selection of the site where these pieces of data will be gathered and the task will be executed and (iii) the selection of the paths the datasets will follow in order to be concurrently transferred to the data consolidating site.Furthermore, the delay required for transferring the output data files to the originating user (or to a site specified by him) should also be accounted for.In most cases the task's required datasets will not be located into a single site, and a data consolidation operation is therefore required.Generally, a number of algorithms or policies can be used for solving these three subproblems either separately or jointly.Moreover, the order in which these sub-problems are handled may be different, while the performance optimization criteria used may also vary.The algorithms or policies for solving these sub-problems compromise a DC scheme.We will present a number of DC schemes.Some consider only the computational or only the communication requirements of the tasks, while others consider both kinds of requirements.We will also describe DC schemes, which are based on Minimum Spanning Trees (MST) that route concurrently the datasets so as to reduce the congestion that may appear in the future, due to these transfers.Our results brace our belief that DC is an important problem that www.intechopen.comAdvances in Grid Computing 96 needs to be addressed in the design of Grids networks, and can lead, if performed efficiently, to significant benefits in terms of task delay, network load and other performance parameters.Information aggregation relates to the summarization of resource information collected in a Grid Network and provided to the resource manager in order for it to make scheduling decisions.Resource-related information size and dynamicity grows rapidly with the size of the Grid, making the aggregation and use of this massive amount of information a challenge for the resource management system.In addition, as computation and storage tasks are conducted increasingly non-locally and with finer degrees of granularity, the flow of information among different systems and across multiple domains will increase.Information aggregation techniques are important in order to reduce the amount of information exchanged and the frequency of these exchanges, while at the same time maximizing its value to the Grid resource manager or to any other desired consumer of the information.An additional motivation for performing information aggregation is confidentiality and interoperability, in the sense that as more resources or domains of resources participate in the Grid, it is often desirable to keep sensitive and detailed resource information private, while resources are still being publicly available for use.For example, it may soon become necessary for the interoperability of the various cloud computing services (e.g., Amazon EC2 and S3, Microsoft Azure) that the large quantity of resource-related information is efficiently abstracted, before it is provided to the task scheduler.In this way, the task scheduler will be able to use efficiently and transparently the resources, without requiring services to publish in detail their resources characteristics.In any case, the key to information aggregation is the degree to which the summarized information helps the scheduler make efficient use of the resources, while coping with the dynamics of the Grid and the varying requirements of the users.We will describe a number of information aggregation techniques, including single point and intra-domain aggregation and we will define appropriate grid-specific domination relations and operators for aggregating static and dynamic resource information.We measure the quality of an aggregation scheme both by its effects on the efficiency of the scheduler's decisions and also by the reduction it brings on the total of resource information.Our simulation experiments demonstrate that the proposed schemes achieve significant information reduction, either in the amount of information exchanged, or in the frequency of the updates, while at the same time maintaining most of the value of the original information.The remainder of the paper is organized as follows.In Section 2 we report on previous work.In Section 3 we formulate and analyze the Data Consolidation (DC) problem, proposing a number of DC schemes.In Section 4 we formulate the information aggregation problem and propose several information aggregation techniques.In Section 5 we present the simulation environment and the performance results obtained for the proposed schemes and techniques.Finally, conclusions are presented in Section 6. Previous workThe Data Consolidation (DC) problem involves task scheduling, data management and routing issues.Usually these issues are handled separately in the corresponding research papers.There are several studies that propose algorithms for assigning tasks to the available resources in a grid network (Krauter et al., 2002).A usual data management operation in grids is data migration, that is, the movement of data between resources.The effects of data www.intechopen.com