Parallel processing of XML databases
Ghassan Z. Qadah · 2006
Beowulf cluster is a name given to a high performance, low-cost parallel computer system made of commodity hardware and software components. It consists of a number of processing nodes, interconnected via a switch. The extensible markup language (XML) data model, on the other hand, has recently gained huge popularity because of its ability to represent a wide variety of structured (tabular-like) and semi-structured (textual-like) data. Several query languages have been proposed for the XML data model, the most-widely known is XQuery. This paper reviews the XML data model and its query language within the context of cluster/parallel computing environment. It examines several techniques for structuring and storing XML data across the different cluster nodes. It develops a number of algorithms suitable for processing a certain class of queries, namely, the containment queries, against the parallel XML database. This paper also shows that one of these algorithms, the one that takes advantage of the parallelism existing between the different documents within the XML database, is outperforming all of the other presented ones