Probabilistic PCA and neural networks in search of representative features for some yeast genome data
Anna M. Bartkowiak, S. Cebrat, Dorota Mackiewicz · 2004
We considered a data matrix with N = 3300 rows (objects, genes) and d = 13 columns representing variables (traits) measured for each object (identifled gene). The 13 characteristics were obtained from so called ’spider-plots’ constructed for each gene. Our goal was to flnd a latent structure in the data and possibly reduce the dimensionality of the data. To achieve this goal we used the methods of ordinary principal components (PC), probabilistic principal components (PPCA) and feed forward neural networks (multi-layer perceptrons). We got some evidence, that H=6 latent variables explain the essential features of the data. Our results are the following: a) First six principal components explain 0.8844 of total variance, however have no interesting interpretation; b) First 6 probabilistic principal components with rotation varimax explain 78.53 % of total variance of the data and have a very interesting interpretation: the set of the primary 12 variables is reduced to 3 double factors, each factor expressed by 2 latent variables; thus we found a meaningful latent structure with a parsimonious representation. c) The multi-layer perceptron with architecture 13{6{13 explains about 88.40 % of total variance, moreover, the matrix of weights, after permuting the columns, yields the same interesting interpretation as the PPCA. Thus a neural network (perceptron) is able to reduce the dimensionality and yield a parsimonious representation of the original variables, similar to that yielded by the PPCA. This { to our knowledge { was not noticed before.