Quick Scanning Huge Dimensional Data

Patrick De Mazière · Lirias · 2013

In statistics one can distinguish three cases: 1) datasets where the number of dimensions is many times larger than the number of samples, 2) datasets where the inverse holds, and 3) both numbers are quite equal (and possible also quite high). Here we focus on the first case, albeit that datasets of the third case might benefit as well from the techniques explained during this lecture. This lecture will start with a brief sketch of two scientific areas, the speaker was actively involved in, that rely heavily on (huge) high-dimensional data sets to find relationships within the data: functional Magnetic Resonance Imaging (fMRI) and semantic text mining. Once the audience has become familiar with some of the properties of these data sets, an overview of (visual) statistical methods is given to explore such data sets in a quick and/or fast manner. Indeed, this kind of datasets is common to many other scientific areas such as aeronomy, econometry, weather, genomics, OCR, …. Consequently, some (basic) exploration methods will come in very useful. In addition, the speaker will also briefly outline/show the benefits of using the High Performance Computing infrastructure (HPC @ KU Leuven, HPC @ VSC) as a way to considerably speed up analyses and/or visualisation. Methods being (briefly) discussed in this lecture are, amongst others, parallel coordinates (& HPC), use of principal components, support vector machines & HPC in the framework of “pure” visualisation. With respect to dimensionality reduction techniques, we discuss some based on distance preservation (projection on principal components – PCA or multidimensional scaling - MDS) and others based on topology preservation (self-organising maps – SOM). Also the use of confusion/correlation matrices & HPC and t-Distributed Stochastic Neighbour Embedding (t-SNE) methods are briefly discussed, not to forget common terms like “the curse of dimensionality”.

Read the paper · More papers on PaperTik