Large datasets: Segmentation, feature extraction, and compression

TN (United States) Oak Ridge National Lab., D Downing, USDOE, Washington, DC (United States), V Fedorov, William F. Lawkins, M Morris, G Ostrouchov · 1996

Large data sets with more than several mission multivariate observations (tens of megabytes or gigabytes of stored information) are difficult or impossible to analyze with traditional software. The amount of output which must be scanned quickly dilutes the ability of the investigator to confidently identify all the meaningful patterns and trends which may be present. The purpose of this project is to develop both a theoretical foundation and a collection of tools for automated feature extraction that can be easily customized to specific applications. Cluster analysis techniques are applied as a final step in the feature extraction process, which helps make data surveying simple and effective.

Read the paper · More papers on PaperTik