A Novel Architecture for Grid Information Systems
Zoltán Balaton, Gábor Gombás, Z. Nemeth · 2003
Grids facilitate large-scale distributed resource sharing.In such an environment a priori knowledge is not availabledue to the diversity of resources and their dynamics. In-formation in the grid ranges from static to highly dynamic.We focus on information required for resource brokering,i.e. how appropriate resources for applications can be foundand what requirements it poses for an information system.One of the representative grid information systems is theGlobus Metacomputing Directory Service (MDS) currentlyin its second incarnation called MDS-2 [1]. It performs bet-ter than its predecessor with respect to scalability and per-formance, yet it still may fail to satisfy every needs of aresource broker. In MDS-2 Grid Information Index Ser-vice (GIIS) servers (specialised indexes of the available re-sources) are used to provide a central repository for clientsto facilitate searches. They collect, cache and manage in-formation from resources belonging to a so called virtualorganisation. It works well in small virtual organisationsbut it is not scalable for thousands of resources.One of the reasons is that MDS-2 aims to provide a uni-form information and monitoring system, i.e. handles staticand dynamic data in the same way. Since frequently chang-ing data becomes stale quickly, the tree of the GIIS serversforminga cache chain cannot be tall. Also as a consequenceof the pull data delivery model used in MDS-2 the amountofdata movedis proportionalto the numberof queries. Thislimits scalability and efficiency.Furthermore, MDS-2 uses the LDAPv3 [4] protocolwhere information is organised as a hierarchical tree calledDirectory Information Tree (DIT) defined by the distin-guished names (DN) of the entries. Although the LDAPquery language permits searches based on any properties ofthe entries, a query that does not match the hierarchy of thedistributed LDAP database can be very inefficient. This isbecause the DIT is distributed along the hierarchy definedby the DNs and one can only restrict the set of servers tobe searched in terms of this hierarchy (with the base andscope parameters in the query). Thus, most queries whichdo not match the hierarchy predefined by the DNs of en-tries result in querying every server storing the distributedLDAP database. Therefore, the distributed LDAP databasecan only answer certain searches efficiently that match thepredefined hierarchy.The solution proposed in this paper tries to tackle withthese problems. Key points are introduced in the following.