Building XML statistics for the hidden web
Ashraf Aboulnaga, Jeffrey F. Naughton · 2003
There is currently a lot of interest in developing Internet query processors that can pose elaborate queries on XML data on the Web. Such query processors can query data sources that have static XML files, but they should also be able to query "hidden Web" data sources that export an XML view of data stored in a database. To optimize queries that involve these hidden Web data sources, we need to have XML statistics that can be used to estimate the selectivity of queries posed to these sources. Since we can only access the data at a hidden Web data source by issuing queries, we need to develop on-line XML statistics that are built by observing queries to a hidden Web data source and their result sizes.