XML keyword search based on maximum repetitive unit

Desheng Wang, Guiquan Liu, Qiming Luo, Enhong Chen · 2011

As XML becomes the standard for data representation and exchange, effective and efficient methods for XML data retrieval have become increasingly important. In practice, XML documents tend to have a shallow and wide structure, and contain a large number of duplicate units. According to these characteristics, we propose a novel algorithm for keyword search in XML documents based on maximum repetitive unit. The basic idea of the algorithm is as follows. Firstly, extract the duplicate structures of XML documents as repetitive units. Then find out which units contain all the query keywords. The results returned are a number of repetitive units associated with the query. The experiments show that the algorithm is scalable, efficient and able to obtain query results with good semantic integrity.

Read the paper · More papers on PaperTik