A compression technique for large statistical data-bases
Susan J. Eggers, Frank Olken, Arie Shoshani · eScholarship (California Digital Library) · 1981
In this paper we explore the compression of large statistical databases and propose techniques for organizing the compressed data.such that the time required to access the data is logarithmic.The techniques exploit special characteristics of statistical databases.namely.variation in the space required for the natural encoding of integer attributes.a prevalence of a few repeating values or constants.and the clustering of both data of the same length and constants in long.separate series.Our techniques are variations of run-length encoding.in which modified run-lengths for the series are extracted from the data stream and stored in a header.which is used to form the base level of a Btree index into the database.The run-lengths are cumulative.and therefore the access time of the data is logarithmic in the size of the header.We discuss the details of the compression scheme and its implementation.present several special cases and give an analysis of the relative performance of the various versions.