Heterogeneous databases—towards a federation of autonomous systems
Marek Rusinkiewicz · Fall joint computer conference · 1987
Heterogenous database systems have recently begun to emerge because the progress in computer networking made their development technologically feasible and, what is more important, in response to an authentic demand from the users of information systems.In any functioning large organization, be it business or government, there are multiple databases in use, supporting their own applications and end-users. These databases may have different scopes, different data models and query languages, etc. Attempts to eliminate or suppress this diversity are usually not successful. Large “comprehensive” and “integrated” database systems built this way frequently become white elephants, irrelevant to the needs of end-users and largely ignored by them.It appears, that the concepts of “global schema” (and its distributed equivalent of “integrated schema”) on which such approaches are based represent an utopian idea that all data resources, for all users can be managed and maintained by a central authority (Database Administrator) who “knows better” what are the needs of the users. The totalitarian overtones of this approach (typical for most utopian ideas) are easily detectable. On the other hand, a consensus seems to exist among users that many applications exist which require access to data from multiple databases. How can such requirements be met?A possible practical approach is to create loose federation(s) of database systems which can cooperate in fulfilling users requests for information. Under this approach the database systems may exchange the (meta)-information about these parts of their data content that they decided to contribute to the global pool of information. If a request for an access to data is made to a local database system it can then locate the data in another system and forward the request to the owner of the data. This process can be wholly or partially transparent to the end-user. However, at each time the local systems exercise the full control over their data, thus preserving their autonomy.Of course, to implement such systems we need to solve number of problems pertaining to data architecture and system architecture. The first group of problems deals with possible incompatibilities among databases participating in a federation. For example, the member databases may have different user interfaces, their schemas may be inconsistent, domain definitions for attributes may be different, the data themselves may be redundant or inconsistent. The second, equally important, group of problems is related to the transaction management in heterogeneous systems. Most of the prototype implementations described in the literature allow only retrieval operations in a heterogeneous environment, since the updates present serious problems in the areas of concurrency control, logging, security, etc. Only when these and other related problems will be solved, the necessary conditions for creating “true heterogeneous databases” will be satisfied.Thus, although we can not fully answer the question WHEN the true heterogeneous systems will become available we may attempt to suggest HOW this can be accomplished. The key to the development of heterogeneous databases is in the evolution of existing systems in an attempt to satisfy real needs of end-users. The inevitable partial loss of local autonomy must be clearly compensated by the expected gains in accessibility to data.