ADAPTIVE QUERY PROCESSING FOR DATA AGGREGATION: MINING, USING AND MAINTAINING SOURCE STATISTICS

Jianchun Fan · 2006

Most data integration systems focus on “data aggregation ” applications, where individual data sources all export fragments of a single relation. Given a query, the primary query processing objective is to select the appropriate subset of sources to optimize conflicting user preferences. We develop an adaptive data aggregation framework to effectively gather and maintain source statistics and use them to support multi-objective source selection. We leverage the existing techniques in learning source coverage/overlap statistics, and develop novel approaches to collect query sensitive source density and latency statistics. These statistics can be used to optimize the coverage, density and latency objectives during source selection. We present a joint optimization model that supports a spectrum of trade-offs between these objectives. We also present efficient online learning techniques to incrementally maintain source statistics.

Read the paper · More papers on PaperTik