Query processing and indexing techniques on semi-structured data
Jun Yang, Hao He · 2007
Management of semi-structured data has become very important in recent years, as the amount of such data and the number of sources producing such data continue to grow. Semi-structured data is often represented as trees or graphs, with text labels on edges and/or nodes. Unlike traditional relational data that conforms to rigid, predefined schema, semi-structured data is schemaless, self-describing, and richer and more flexible in structure. In addition to documents and Web pages, semi-structured data also comes from an increasing number of applications based on XML-related standards. Scientific applications are major users too, because semi-structured data is well suited to handling complex, heterogeneous, and evolving data standards. Personal information management and social networking applications also produce semi-structured data whose graph structure captures connections among entities. Queries over semi-structured data consider its textual contents as well as structure. Query processing is challenging because of the lack of schema and richness in structure. This dissertation develops a collection of query processing and indexing techniques to support efficient queries over tree- and graph-structured data. Specific contributions include (1) practical index structures for supporting evaluating label path expressions and checking graph reachability, two fundamental primitives for query processing over semi-structured data, and (2) a keyword search system for finding and ranking substructures of interest within text-labeled graph-structured data.