Subject Identification in Topic Maps in Theory and Practice.

Lutz Maicher · Berliner XML Tage · 2004

If Topic Maps should be exchanged in distributed environments a common semantic problem occurs: Do two Topics represent the same Subject? If they describe the same Subject the according Topics have to be merged. Within the Topic Map theory the merging paradigm and the description of Subjects is the main theoretical design criterion. Normally, these methods provided by the standard lead to sufficient results, but only if distributed Topic Map authors share a common vocabulary for Subject description. To solve the arising problems for distributed, autonomous environments we introduce the Subject Identity Measure. The SIM describes how closely related the Subjects of two distributed Topics are. The approach is independent from a shared vocabulary, from a specific natural languages and uses only data which is available inside the according Topic Maps. 1 Problem Subject Identification in distributed Topic Maps A Topic is a binding point for all information concerning one Subject within a Topic Map. A Subject is anything of the real world where an author wants to discourse about within his Topic Map. The main theoretical design criterion of Topic Maps is called “One Topic for one Subject”. This means, if two Topics describe the equal Subject (this decision is supported by equality rules) within the same Topic Map they must be merged to one Topic (this process is defined by merging rules). “One Topic for One Subject” is well defined and applied within the Topic Map standards. Problems occur if the Subjects of Topics aren’t described with a shared vocabulary; especially in distributed environments. In these cases the defined equality rules have strong limitations or fail because Subjects are only defined as identical if they are described with identical strings. We don’t share the optimism, that centralised repositories for Subject description (so called PSI repositories) will be widely adopted. Rather we expect that the sameness of two Subjects can be inferred from the content of its Topics. Therefore we introduce the Subject Identity Measure which describes how closely related the Subjects of two Topics are. The level of the SIM supports humans or machines to decide whether these Topics should be merged or not. In general, we foresee interesting applications for the usage of the SIM. Distributed knowledge management might be the main application [see Cu03, Sc04, and Si04]. Topic Maps are part of the Semantic Web efforts [see Th02] and are translatable into RDF or OWL [discussed in Pepp, Gars, and PS03]. This enables the reuse of the SIM approach in a variety of Semantic Web applications. At least, the SIM approach can be used for the integration of unstructured and structured information in business processes. In this paper we are making the following contributions: « We discuss the arising problems of Topic Maps’ central theoretical criterion „One Topic for one Subject“ (see section 2). « We describe the Subject Identity Measure approach to address these problems. Additionally, we assess its quality yielded for a testbed in brief (see section 3). « We sketch the challenges of further research (see section 4). 2 Subject Identification in the Topic Map Theory The main theoretical design criterion of Topic Maps is called “One Topic for one Subject”. In order to understand this criterion, we need to explain the notions of Topic, Subject and their relationship. A Topic is “a symbol used within a topic map to represent some subject, about which the creator of the topic map wishes to make statements” [TMDM]. A Subject is “anything whatsoever, regardless of whether it exists or has any other specific characteristics, about which anything whatsoever may be asserted by any means whatsoever” [TMDM]. Shortly, a Topic describes a Subject (which is anything on which the creator of a Topic Map chooses to discourse) from the perception of the current Topic Map. This implies, within each Topic the Subject must be declared. While declaring Subjects, important philosophical questions arise: What is identifiable? What constitutes the boundaries of a thing in respect to its identity? Can identity evolve in time? Is identity situational or relative? How must properties of a thing change to alter its identity? What about versions and copies? These questions [discussed in detail in Ke78, Ke03] show the limits of pure naming approaches because they hardly handle indefiniteness, openness and ambiguity [see FLGD87]. But how a Topic can declare its Subject? Within the Topic Map Data Model (TMDM) two means are implemented which are more or less pure naming approaches: The Subject Locator is used whenever the Subject of the Topic is an addressable information resource. In this case, the URI of this resource is used as a Subject Locator. The URI names the Subject. Because Subjects can be anything (not only addressable resources) a Topic can declare its Subject with the help of a Subject Indicator, too. A Subject Indicator is an information resource which describes the Subject. The URI (which names the Subject) of this information resource is called Subject Identifier. To obtain “One Topic for one Subject”, two Topics having the same Subject Locator or a pair of identical Subject Identifiers have to be merged. These rules work well if all authors of Topic Maps have made agreements about a shared vocabulary for Subject naming. These agreements are called Published Subject Indicators (PSI) [Oasis]. PSIs are published (but not necessarily public) descriptions of Subjects which should be reused by as much Topic Map authors as possible to obtain a broad interoperability of Topic Maps. Examples in the literature which discuss merging of distributed Topic Maps (or Topic Maps and RDF documents) exclusively use PSIs [see CPV03, Gr02, Sc04]. This is due to the absence of solutions for open vocabularies. However, in distributed environments with a high autonomy, the mechanism of PSIs has its shortcomings. PSIs are only used if they are visible to the regarding Topic Map authors. Additionally, PSIs are faced with the philosophical problems of naming approaches discussed above. In contrast to the naming approach we follow up an description approach. We assume that a Subject is indirectly determined by the content of its Topic. We don’t name or stringently delimit a Subject, we only decide whether two Subjects are quite identical. The level of this “identity” supports humans or machines to decide, whether these Topics should be merged. If they chose merging, these Topics get an identical Subject Identifier to apply the merging inside the Topic Map standards. 3 The Subject Identity Measure (SIM) Approach But “Merging beyond the minimal rules [defined in the TMDM] is freely allowed. Most commonly, this will be done by inferring the subject of the topics from their characteristics.” [TMDM]. Therefore, we propose a Subject Identity Measure. The SIM describes how closely related the Subjects of two Topics are. If the SIM is 1 the regarding Topics definitely represent the same Subject (according to the rules defined in the TMDM). If 0, the regarding Topics definitely represent different Subjects. All values between 0 and 1 support the decision whether two Topics represent the same Subject. Whenever two Topic Maps meet, the SIM Approach performs the following steps: 1. Calculation. The SIMs for TopicNames, Occurrences, and Subject Indicators for each pair of Topics must be calculated. 2. Filtering. According to different thresholds and coefficients, the overall SIM will be calculated. For each Topic, a suitable counterpart in the other Topic Map will be chosen (but only if there is a SIM greater than 0). For calculation only data inside each Topic is used, whereby structural (types, associations etc.) or external (content of information resources referenced from the Topic Map) information is left aside. We need similarity measures for URIs (for Subject Indicators, VariantNames and OccurrenceLocators) and strings (all TopicNames and OccurrenceData). For the calculation of the similarity of a pair of two strings (S1,S2) we used a language and context independent measure c(S1,S2)->[0,1]. For more detail and the discussion of the calculation of the URI similarity we refer to [MW04]. In general, the approach inspects each possible pair of Topics (T1,T2) where T1 and T2 belongs to two different Topic Maps. For each pair (T1,T2) the measures SIM.Names and SIM.Occurrences are calculated as follows: 1. “Fillet” each Topic for Names. Take all property values from the property “value” of all Topic Name Items and Variant Items of T1 and store them in a set Nam1. To get Nam2 do the same for T2. 2. “Fillet” each Topic for Occurrences. Take all property values from the property “value” of all Occurrence Items of T1 and store them in a set Occ1. Do the same for T2. 3. Calculate SIM.Names. If |Nam1| < |Nam2|, and |Nam1|=m then:

Read the paper · More papers on PaperTik