A Study on Effective and Efficient XML Element Retrieval Considering Document Updates
Atsushi Keyaki · Institutional Repositories DataBase (IRDB) · 2014
Extensible Markup Language (XML) is a markup language for structured documents that has become the de facto format for data exchange.The data formats of many applications, such as Wikipedia articles and office documents, are defined as XML formats.A large number of XML documents are available on the Web, and we expect this trend to continue in the future.Search system users have to find out information they need by themselves, because most of existing search systems return a list of documents as search results.It takes a large effort to find it out from long-length documents.In addition, there is a possibility that users cannot find useful information nevertheless they spent long time on the documents.This information seeking process is much cost in information retrieval.On the other hand, users of XML element search systems need not to find out information they need by themselves because XML element search systems return a list of elements which satisfy users' information need.Therefore, the framework of an XML element search system can reduce the cost in information retrieval, which is the reason why XML element retrieval techniques are useful and worth working on.There are two main streams for researches of XML element retrieval techniques, i.e., 1) attaining effective search for satisfying accurate information retrieval, and 2) attaining efficient search for fast query processing.In order to satisfy 1), we proposed a scoring method to identify informative XML elements and a reconstruction method of search results to identify the most appropriate