Special Issue: Selection of Best Papers of the VLDB Data Management in Grids Workshop (VLDB DMG 2006)
Jean‐Marc Pierson, Lionel Brunie · Concurrency and Computation Practice and Experience · 2007
Grid computing exists now at a widely distributed scale.After 10 years of internationally combined efforts to develop from a vision to existing middlewares, we face now a large number of applications being deployed and taking benefit from the Grids.Almost all these applications handle data, at different semantic levels, from raw data coming from sensors in particle physics to rich data in the healthgrids.Most of the time, the procedure is ad hoc: The need to handle these data has been mainly seen as a constraint, and little effort has been put in their smart management in the early stages of the grid evolution.Data were present in raw files, or in databases, may be distributed databases, but with little concerns from the application developer who focused (and that is normal) on the core development of the process of the data.In the close past, the Grid community has been developing specialized services to handle data in a simpler and more smart way: OGSA-DAI allows for instance to access several data sources with a common programming interface, giving the developer also the possibility to add treatment on the data retrieved, to anonymize it, to cache it, or to replicate it.Efforts have been put to access and process the data, not in planning optimization or distributed balanced queries, which are core distributed databases services.Leading databases companies have publicized grid-aware databases that mainly are 'old parallel databases' in clusters, not taking much into account the challenges of the grid environment.The database community has been investigating for a long time issues related to the distribution of the data sources, data queries, query plan optimization in parallel systems but has not been really involved up to know in the grid community and development.Among the challenges rising in these environments, we can cite, among others: dynamicity, reliability, security, data availability and transport, data indexing, search and access, etc.These are existing distributed databases challenges revisited with respect to the grid paradigm.This special issue of 'Concurrency and Computation: Practice and Experience' includes a selection of revised papers presented at the VLDB Data Management in Grids workshop, Seoul, Korea, held on 11th September 2006.They illustrate some of the challenges together with current trends of solutions.The 2006 edition of the workshop is the second in a raw, after the success of the 1st edition in Trondheim in 2005.The idea of the workshop collocated with one of the most important conference in databases (VLDB: Very Large Data Bases) is to bring together the experts from the communities to meet, discuss and argue the above stated challenges.In that sense, the objective of the workshop