Classification Rules for Semi-Structured Data.
David Michaeli, Werner Nutt, Yehoshua Sagiv · 1997
this paper, we want to exploit this similarity to apply reasoning techniques from description logics to semistructured data. Semistructured models are intended to capture data that are not intentionally structured, that are structured heterogeneously, or that evolve so quickly that the changes cannot be reflected in the structure. A typical example is the World-Wide Web with its HTML pages, text files, bibliographies etc. Another example are biological databases that are often realized as files, but users want to access them in an integrated fashion. By their very nature, semistructured data do not come with a conceptual schema. However, adding to them a rich conceptual model is beneficial, since it would make them more accessible to users. This is especially important, since semistructured data are often accessed in an "explorative" or "browse mode," i.e., users not only query the data to find a particular piece of information, but also pose queries with the goal of having a better understanding of what information is available. We propose in this paper to build a layer of classes on top of a semistructured data model that is an abstraction of the OEM model. The classes are defined by rules and populated by computing a minimal fixpoint. We consider two semantics for such rules, which we call "strong" and "weak" semantics. Under the strong semantics, classes are only populated by the rules, while under weak semantics users are allowed to arbitrarily