Efficient indexing techniques for XML data

Dao Dinh Kha · Institutional Repositories DataBase (IRDB) · 2004

In today's just-in-time information paradigm, the ability to communicate efficiently is vital.In this context, Extensible Markup Language (XML) has been invented as a standard for data exchange among a variety of data sources and applications.Since its invention, the popularity of the XML has been dramatically increased thanks to the ability of XML to provide a simple but standardized, extensible means to express semantic information within documents.This property makes XML possible to address the shortcomings of existing markup languages and become a key technology that facilitates information exchange for science, technology, and industries, as well as many aspects of society.XML also has been selected as the data exchange standard in newly arising business domains, such as e-business.The semantics of XML data is richer than the one of relational data by including both of its content and metadata, the information that describes the structure of the data.Therefore, XML requires new techniques that are different from the relational databases technologies for data management.The aim of this thesis is to design efficient indexing techniques for XML data.The structural characteristics of XML data raises several challenges to design of indexing structures.In this thesis, we investigate three critical issues of XML data management as follows.The first issue is to cope with intensively content-updated XML data.Among several methods of storing XML documents, a straightforward yet efficient method is to store a string representation of the XML document.An XML node is usually represented by a region coordinate, which is a pair of integers expressing the start and end positions of the substring corresponding to the node.This approach, however, has the drawback that a change of a node's region coordinate causes change of the * Doctor's Thesis,

Read the paper · More papers on PaperTik