The Organization of Information in a Statistical Office
Tjalling Gelsema · Journal of Official Statistics · 2012
A theoretical framework for statistical data and metadata is presented using well-known techniques from computer science. It is shown that many kinds of statistical information are best represented by a function. The relationships between different pieces of statistical information, such as the dependency between aggregated data and the underlying microdata, are explained using transformations of functions. A modest set of such transformations is considered and the identities that hold between them are shown. Thus they describe a class of algebras, and, according to the theory of initial algebra semantics, the initial algebra in the class is the natural candidate for recording metadata.