Identity and Versions for Complex Objects.

George P. Copeland, Setrag N. Khoshafian · 1986

Identity is that property of an object that distinguishes each object from all others. Identity has been investigated almost independently in general-purpose programming languages and database languages. Its importance is growing as these two environments evolve and merge. Historical versions of persistent data is a natural and useful feature for a large number of applications. Identity is even more important when retrievals may span multiple versions of a single object, since a mechanism is needed to preserve the fact that the various versions are of the same object. There are at least two dimensions involved in the support of identity, the representation dimension and the temporal dimension. The representation dimension distinguishes languages based on whether they represent the identity of an object by its value (e.g., identifying employees by social security number), by a user-defined name (e.g., variable names, user defined file names, etc.), or built into the language (e.g., Smalltalk-80). A language providing a stronger notion of identity in this dimension must maintain its representation of identity during updates, use identity in the semantics of its operators, and provide operators to manipulate identity. The temporal dimension distinguishes languages based on whether they preserve their representation of identity within a single program or transaction, between transactions, or between structural reorganizations. An example of structural reorganization is schema reorganization in databases. A language providing stronger identity in the temporal dimension must employ more robust implementation techniques to preserve its representation of identity. The thesis of this abstract is that database conceptual languages should support the built in notion of identity and provide for the possibility of structural reorganizations. There are very few conceptual database models (in fact one we know of) that support this very strong notion of identity. Most real-world organizations deal with histories of objects, but they have little support from existing systems to help them in modeling and retrieving historical data. Strong support of identity in the temporal dimension is even more important for temporal data models, because a single retrieval may involve multiple historical versions of a single object. Such support requires the database system to provide a continuous and consistent notion of identity throughout the life of each object, independently of any descriptive data or structure which is user modifiable. This identity is the common thread that ties together these historical versions of an object. Most of the current techniques for implementing object identity (such as as physical addresses, or identifier keys (i.e. groups of attributes which constitute the primary key) etc.) lack either location independence or data (i.e. object content) independence. The strong notion of identity is most easily supported using surrogates, which are system-generated, globally unique identifiers, completely independent of any physical location.

Read the paper · More papers on PaperTik