Supporting statistics in extensible databases: a case study
A. Segev, A. Chatterjee · 2002
This paper presents a framework for supporting regression in extensible databases. This work is motivated by an actual case study that required probabilistic record matching using regression be supported in a database. We discuss how the regression function can be implemented in a database using rules and functions. We present some indexing and caching ideas for efficient data processing for retrieval-intensive applications. We also describe how the regression coefficients can be incrementally updated by materializing some of the intermediate statistical results. Finally, we illustrate how sampling techniques can be used to generate the data sets for regression and highlight some of the performance issues in this regard.>