Management and display of data analysis environments for large data sets

Robert Alexander Burnett, Paula J. Cowley, James J. Thomas · Statistical and Scientific Database Management · 1983

Data analysis is typically an iterative process in which the choice of the next analysis operation is largely determined by the results of previous operations on the set. With large sets, many analysis paths may be explored before meaningful results are obtained. Along eachpath, the analyst creates a sequence of data analysis environments, each environment being a frame or snapshop of the set and associated descriptions, conditions, models, and analysis results. The analysis environment may be changed incrementally through temprary modifications, subsets, samples, or statistical operations; or, the analyst may wish to restore the conditions of a previous environment as a starting point fram which a new analysis path can be generated. Existing analysis systems, however, lack facilities to maintain, save, or restore all of the components required to completely describe or reconstruct a analysis environment.This paper describes ongoing research at Pacific Northwest Laboratory (PNL) in management and display techniques for multiple analysis environments. Specifically, research is being conducted in four major areas: (1) the development of a model of the analysis process incorporating the concepts of analysis environments; (2) the design and use of modification definitions (differential files) to represent multiple versions of a large base; (3) the use of dictionaries/directories to manage, describe, and control multiple analysis environments; and (4) the application of graphical display and interaction techniques to the examination and selection of analysis environments. The results of these research efforts will be integrated to provide a new dimension in interactive analysis.

Read the paper · More papers on PaperTik