Comparable datasets in performance benchmarking
David M. Steier · 1995
this paper) depends on the task for which the information is being gathered, the target collection of objects to report on, and the data available about each object. The purpose of this workshop paper is to highlight the importance of this problem in gathering information from heterogeneous sources. and to present some detail about a case study encountered in practice while doing a performance benchmarking study. Aspects of producing compsets have been studied in the database literature within the area of schema integration for heterogeneous databases [Batini et al., 1986], because of the shared concern for semantic comparability at the schematic level. For example, the theory of semantic values developed by Sciore et al. [1994] seems like a promising approach to computing comparable datasets because of the explicit representation of contextual information for each value. We discuss a number of issues involved in using contextual information in this way.