Reduction of dimensionality: inferential aspects of descriptive methods

W. J. Krzanowski · 2000

Abstract We now return to some of the techniques already described in Part I, and consider ways in which their purely descriptive use can be underpinned by statistical theory. Broadly speaking, the majority of these techniques comprise two main operations: finding a suitable small-dimensional space in which the sample can be embedded, and then inspecting the representation of sample members in this space for any evidence of pattern. Both of these operations carry associated statistical questions, which were hardly even alluded to earlier. In the first place, what is the ‘correct’ number of dimensions in which the sample is located? Given an r-dimensional representation of a p-variate sample, is this only an approximation, or have we identified the ·true’ dimensionality? Secondly, having obtained a small-dimensional representation, what effect will sampling variation have on the positions of the various entities in this space? For example, how far apart must be two points (or vectors, or curves) before we can safely assume that the entities they represent are genuinely different? In order to answer any of these questions we cannot rely simply on our geometrical model, but must also assume a statistical model for our initial multivariate sample. Hence consideration of these matters belongs properly to this part of the book. However, at this point we will break slightly with our problem-oriented philosophy. Since we have already devoted much space to the introduction and description of the various techniques earlier, it seems more natural to switch (in this chapter only) to a technique-oriented discussion, in which all aspects associated with statistical inference are summarized for each technique in turn.

Read the paper · More papers on PaperTik