08421 Working Group: Lineage/Provenance
Anish Das Sarma, Amol Deshpande, Thomas Hubauer, Ihab F. Ilyas, Birgitta König‐Ries, Matthias Renz, Martin Theobald · 2008
Abstract. The following summary tries to capture a collection of state-of-the-art techniques and challenges for future work on lineage management in uncertain and probabilistic databases that we discussed in our working group. It was one half of a larger committee that we had initially formed, which then got split into two groups—one focusing on lineage as a means of explanation of data, and one focusing more on lineage usage in probabilistic databases (see also the “Explanation” working group report for more details on the first subgroup). 1 Why lineage? Users want lineage as a basic functionality of the database or information system they use. Explicitly tracing lineage information along with the actual results of queries may help users to better understand complex workflows or to analyze results in huge data warehouses that are often derived from complicated queries and many levels of materialized views. Lineage, as a form of annotation or metadata over the actual data, can provide this information without the need to rerun the entire query or workflow and thus help to trace the origin and/or explain the derivation of individual results. Tracing