Special Issue: Selection of Best Papers of the VLDB Data Management in Grids Workshop (VLDB DMG 2007)
Jean‐Marc Pierson, Harald Kosch · Concurrency and Computation Practice and Experience · 2008
Grid computing has greatly matured since the last decade. It exists now at a widely distributed scale. After 10 years of internationally combined efforts to develop from a vision to existing middlewares, we now face a large number of applications being deployed and taking benefit from the Grids. The need to handle data properly came with the applications, managing data at different semantic levels, from raw data coming from sensors in particle physics to rich data in Healthgrids. Most of the time, data management was ad hoc: The need to handle these data has been mainly seen as a constraint, and little effort has been put in their smart management in the early stages of the grid evolution. Data were present in raw files, or in databases, maybe distributed databases, but with little concerns from the application developer who focused on the core development of the process of the data. In the close past, the Grid community has been developing specialized services to handle data in a simpler and more smart way: Pieces of middlewares have been developed, for instance, to access several data sources with a common programming interface, giving the developer also the possibility of adding treatment on the data retrieved, anonymizing it, caching it, or replicating it. Works on data caching, replication, data integration, and security, to name but a few, have been seen. Efforts have been put to access and process the data, not in planning optimization or distributed balanced queries, which are core distributed database services. The database community has been investigating for a long time the issues related to the distribution of the data sources, data queries, query plan optimization in parallel systems, etc. Among the challenges rising in grid environments, we can cite, among others: dynamicity, reliability, security, data availability and transport, data indexing, search and access, autonomics, etc. These are existing distributed database challenges revisited with respect to the grid paradigm. This special issue of ‘Concurrency and Computation: Practice and Experience’ includes a selection of revised papers presented at the VLDB Data Management in Grids workshop, Vienna, Austria, held on 23rd September 2007. The 2007 edition of the workshop was the third in a raw, after the success of the 1st edition in Trondheim in 2005 and the second in Seoul in 2006. The idea of the workshop collocated with one of the most important conferences in databases (VLDB: Very Large Data Bases) is to bring together the experts from the database and grid communities, discuss, and argue the above-stated challenges. In that sense, the objective of the workshop is fulfilled, with several interesting emerging discussions during the event that may be reflected in this Special Issue. Efficient grid data management and integration is achieved only through a perfect teamwork of all levels in grid databases, i.e. the conceptual level, the middleware level for service integration, and finally the low level for data processing and redistribution. The following papers give a good overview of methods and implementations for efficient data management and integration at these levels. The paper authored by Göres and Dessloch 1 presents the PALADIN Integration Framework that allows the integrated representation of arbitrary data, metadata, and operations. On top of it different integration services can be more efficiently built, as their describing metadata have already been integrated. Rabl et al. 2 deal with the low-level aspects of data integration and management. In dynamic grid environments, automatic data management is one key issue for good performance. The paper proposes a new algorithm for load balancing which adapts the data management to the automatic scaling in the number of nodes used. The second and third papers go ‘middleware’: Kotowski et al. 3 present the GParGRES Grid database, which originally implements a middleware solution among cluster databases to enable data grid services. The architecture scales perfectly. Bassi et al. 4 present a multi-layer solution to tackle the outstanding data to be produced and managed in the forthcoming years. Their solution named Intelligent Network Caching Architecture (INCA) combines new networking entities, dedicated transport protocols, peer-to-peer middleware placement, and access functions in a comprehensive framework. Two papers dealt with replication or file management and access. In 5, Akal et al. presented a new replication scheme for grid databases following a publish/subscribe model. Principally, a middleware over the different grid nodes thus allows one to dynamically control the creation and maintenance of replicas. In 6 Hupfeld et al. analyze how an object-based file system can be adapted to Grid environment to provide an abstraction layer to the distributed file management procedure. Then they describe the XtreemFS architecture and explain how it can serve as an improvement of current data management in Grids. Santos and Koblitz present in 7 some security issues related to the management of replicated metadata catalog in the EGEE European project. They explain how security is handled in the AMGA metadata catalog, and its link and use of common security mechanisms in EGEE and Globus-based Grids (VOMS, GSI). In particular, the authors discuss the authorization policies with replicated data, illustrated with a usecase in the healthgrid field. The last paper in this special issue is authored by Norman W. Paton, invited speaker at the VLDB DMG'07 workshop. In 8, he provides an interesting position paper on the state of the art on functions already automated in the current approaches to adaptive data management (for instance, in database administration, query processing, data integration). He argues about some limitations in the current approaches of autonomic computing for data management (predictability, methodology, composability, and semantics) and highlights new challenging researches to handle automation in all aspects related to data management. These eight articles reflect the common trends of research in Grid-related projects linked with data management. We have seen a particular interest this year on adaptability and data management to make the data management in grids more sensitive to the changes in the infrastructure. We hope you will enjoy this selection of papers as much as we enjoyed them during the workshop and during the last reviewing phase. We would like to take the opportunities to thank all our Program Committee members who helped a lot during the whole process as well as G. Fox for his help to make this special issue possible. Finally, we invite you to participate in the 4th edition of the workshop that will be held at VLDB in Auckland in 2008.