Operator-Based Query Progress Estimation

Ralf Hartmut Güting · 2008

Abstract: Recently, research has addressed the problem of estimating progress for long-running data-base queries. The basic idea is to “continuously ” monitor execution to keep track of how much work has been done, and at the same time to collect statistics to arrive at a more and more refined estimate of the total amount of work that is needed. Previous research has generally decomposed the operator tree for the query into pipelines (or “segments”) of non-blocking operators, tried to observe progress per pipeline and then to combine progress measures of the different pipelines into an overall progress measure. It has soon become apparent that pipelines of non-blocking operators are too large units and that it is necessary to define smaller segments (e.g. containing only one join operator). In this paper we take a more radical approach where each operator in a query tree is able to estimate the progress achieved for its subtree based on the progress reported by its children. No global analysis of the query tree is needed, nor is it necessary to determine driver nodes or dominant inputs. Each operator is strictly independent in its progress estimation. Nevertheless progress estimation works fine across block-ing operators and for the whole query tree. The technique lends itself to a simple and clean implementa-tion. It is suitable for extensible database architectures where the set of query processing operators is large and possibly extended at any time. Our implementation allows one to add progress support for operators gradually such that the system runs at any time and reports progress whenever all operators in the query tree support progress. We report a prototypical implementation in the SECONDO extensible database system. Progress estimation now is a standard feature of SECONDO. To our knowledge it is the first freely available DBMS prototype that includes query progress estimation. 1

Read the paper · More papers on PaperTik