Semantic metadata driven data analysis
Sharad Mehrotra, Dawit Yimam Seid · 2007
With the growing availability of diverse and complex data to a wide range of users through the Web, two major information access challenges are becoming increasingly important. The first one is enabling powerful querying, exploratory and analytical operations on unstructured or minimally structured (extracted) data beyond keyword search. The second one is enabling flexible and navigational operations on structured data containing complex relationships. This thesis explores and demonstrates the role semantic metadata plays in addressing both of these challenges. First we consider the problem of developing analytical operators for unstructured data such as text and other media collections. Advances in information extraction and metadata annotation techniques are making it possible to automatically annotate such data with various kinds of taxonomies. Presently, these taxonomies are used just for navigation. We develop several, well-founded analytical operators and algorithms that exploit such taxonomies for novel data summarization and comparative analysis. We then consider the problem of exploring structured datasets that include complex semantic relationships. Presently, the use of semantic metadata to explore structured data is limited to OLAP dimension hierarchies which require a well-crafted data representation. We leverage the approach of multi-faceted (dynamic) navigation used for document collections to develop novel techniques and algorithms for exploration of data that contains complex semantic relationships. We then turn to the problem of efficient query processing on minimally structured data where semantic metadata is uniformly represented along with the data as a semantic graph. We specifically focus on the efficient execution of exact and inexact graph pattern queries on such data. For exact match queries, we develop a complete query algebra that addresses open problems with regard to efficient graph extraction, grouping and aggregation. For inexact match queries, we develop a ranking algorithm that systematically combines both partial structural matches and semantic similarities in order to return semantically-ranked results.