Exploit keyword query semantics and structure of data for effective XML keyword search

Khanh Nguyen, Jinli Cao · 2010

Keyword search is a natural and user-friendly mech-anism for querying XML data in information systems and Web based applications. One of the key tasks is to identify and return meaningful fragments as re-sults, due to the limited expressiveness and the am-biguity of keyword queries. In this paper, we first studied query keyword patterns in order to exploit the user’s search intention behind the input keywords. The outcome of this task is that keywords in the query are classified as required information and search con-ditions (or predicates). In addition, unlike previous work that our work only returns desired fragments as results. Each returned result must satisfy the search conditions rather than simply contain all query key-words. To further prune irrelevant fragments we in-troduce a novel notion called Relevant Lowest Com-mon Ancestor (RLCA) which effectively and precisely captures the meaningful and relevant fragments to the given keyword query. We conducted extensive exper-imental studies to prove the effectiveness of our ap-proach.

Read the paper · More papers on PaperTik