7. Sampling Schemes for a Fixed Dataset

Society for Industrial and Applied Mathematics eBooks · 2005

Since data-mining applications often deal with one or more large, fixed datasets, an important class of sampling schemes consists of those based on subset selection, as noted in Chapter 6. This chapter describes four general classes of subset selection strategies: random subset selection, subset deletion, comparison-based strategies, and systematic approaches that often use auxiliary knowledge about a dataset. Sec. 7.1 describes each of these four approaches in detail and Secs. 7.2 through 7.5 present GSA case studies that illustrate each strategy in turn. Besides their utility in the context of GSA sampling schemes, it is important to note that these subset selection strategies also play an important role in the development of computational algorithms for large datasets, the design of moving-window data characterizations for time-series data, and the stratification of composite datasets for simpler and often much more informative data analyses, relative to “large-scale” analyses of the entire dataset. 7.1 Four general strategies Conceptually, one of the simplest ways of generating fixed-size subsets of a given dataset is random selection, a term whose various possible interpretations are discussed briefly in Sec. 7.1.1. Sometimes, the selected subsets are mutually disjoint (i.e., Si ∩ Sj = ∅ if j ≠ i), but in other cases they are not. This distinction is important since characterizations of overlapping datasets tend to be correlated, an issue examined in Sec. 7.1.2. The degree of overlap is most extreme in the case of deletion strategies, where small subsets are omitted from a larger dataset These strategies can be quite useful in detecting various classes of outliers, an idea discussed in detail in Sec. 7.1.3. The philosophy behind deletion strategies is to make what should be small changes in a dataset and look for unusually large effects. This philosophy is opposite to that of comparison strategies such as the computational negative control idea discussed in Sec. 7.1.4, where we deliberately make what should be large (i.e., “structure-destroying”) changes and examine the consequences: if the changes in computed results are large enough, they provide evidence in support of the hypothesized structure on which our data analysis is based. Finally, Sec. 7.1.5 describes a variety of systematic or partially systematic subset selection strategies that can incorporate explicit prior knowledge, related auxiliary data, or preliminary data characterizations such as cluster analysis results.

Read the paper · More papers on PaperTik