Using Map and Reduce for Querying Distributed XML Data
Lukas Lewandowski · KOPS (University of Konstanz) · 2012
Semi-structured information is often represented in the XML format. Although, a vast amount of appropriate databases exist that are responsible for efficiently storing semistructured data, the vastly growing data demands larger sized databases. Even when the secondary storage is able to store the large amount of data, the execution time of complex queries increases significantly, if no suitable indexes are applicable. This situation is dramatic when short response times are an essential requirement, like in the most real-life database systems. Moreover, when storage limits are reached, the data has to be distributed to ensure availability of the complete data set. To meet this challenge this thesis presents two approaches to improve query evaluation on semistructured and large data through parallelization. First, we analyze Hadoop and its MapReduce framework as candidate for our distributed computations and second, then we present an alternative implementation to cope with this requirements. We introduce three distribution algorithms usable for XML collections, which serve as base for our distribution to a cluster. Furthermore, we present a prototype implementation using a current open source database, named BaseX, which serves as base for our comprehensive query results. iii Acknowledgments I would like to thank my advisors Professor Marc H. Scholl and Professor Marcel Waldvogel, who have supported me with advice and guidance throughout my master thesis. Thank you also for the great possibility to work in both the DBIS and the DISY department and for provisioning a comprehensive workplace and all necessary work materials. Also I would like to thank Dr. Christian Grün and Sebastian Graf for the many helpful