Parallel processing search engine for very large XML data sets

Vu Le Anh, Attila Kiss, Zoltán Vincellér · 2007

In this paper, we study several fundamental problems for building a parallel processing search engine for very large XML data sets. In our model, the data set can be considered as a very large XML file, which is fragmented and stored on different machines. The machines are connected by the high speed local network and the query language is based on regular queries. The following problems are introduced and studied: 1. The general model for the system; 2. The query language; 3. The efficient query processing algorithm; 4. The efficient fragmentation algorithm. All introduced algorithms and solutions bring into play the extensibility and the self-desc ribing of XML data sets, and they also promote the abilities of the parallel systems.

Read the paper · More papers on PaperTik