Discovering Attribute Locality across the Deep Web: an Ordering-Based Approach

Chengkai Li, Kevin Chen–Chuan Chang · 2003

The large number of structured database sources on the Web presents pressing need for information integration at a large scale. How can we enable systematic access to this “deep Web”? We observe that, while au-tonomous sources are seemingly independent, their query schemas often reveal certain cor-relations, such that sources in the same struc-tured domain (e.g., books, cars) tend to share a “locality ” of query attributes. This paper thus develops the notion of attribute locali-ties, which is key for many schema-based in-tegration tasks – such as source clustering and query mediation. Such attribute localities, while very useful, are computational expen-sive to discover. However, our observation further indicates that the localities are often self-revealing, when attributes are linearly or-dered in a certain way, reflecting their con-nectivities. We thus further propose a novel ordering-based approach, which discovers lo-calities by progressive construction, guided by attribute connectivities. Our experimental s-tudy shows that our notion of localities natu-rally capture the inherent structured domains, and our approach effectively discovers such lo-calities, over hundreds of real Web sources. 1

Read the paper · More papers on PaperTik