Improving the performance of lineage tracing in data warehouse
Satyadeep Patnaik, Marshall Meier, Brian Henderson, J. R. Hickman, Brajendra Panda · 1999
The primary use of data warehouse is for decision support.Data warehouse mostly consists of derived data, generally stored as views, rather than raw data, to help in decision making.Sometimes decision-makers want to know the original tuples from which the view has been derived.The purpose of this may be to determine the authenticity of the original data or to bolster the decision-making capabilities.This problem of retrieving the original tuples of a view is termed as data lineage problem or lineage tracing problem.In our work we are suggesting a method, which will perform the task of data lineage with significantly less time.Our method suggests adding an auxiliary column to all tables in the database to store surrogate values for each tuple.Although this method requires extra disc space, the significant saving in response time outperforms this shortcoming.