One Hidden Layer Linear Networks and Canonical Correlations

Zoubin Ghahramani · 2003

This is a short note that relates learning one hidden layer neural networks and the statistical method known as canonical correlations analysis (CCA). It is assumed that the reader is familiar with CCA. [Abstract added: Nov 28, 2003] Consider approximating a function from a p-dimensional input x to a q-dimensional target y, using a linear network ŷ = VW ′x, where the hidden units z = W ′x compute an m-dimensional projection of the inputs. We are interested in the class of linear networks where m < min(p, q), such that the hidden layer constitutes a bottle-neck. It is known that when y = x, the hidden unit representation spans the space of m Principal Components of x (eigenvectors with largest eigenvalue of the x covariance matrix, Σxx)—we provide a generalization of this result. Let the cost C be the expected squared error measured using a Mahalanobis metric based on the output covariance matrix, Σyy, and assume that both inputs and outputs have zero mean, C = 1 2 〈(y − ŷ)′Σ−1 yy (y − ŷ)〉 (1) = 1 2 〈(y − VW ′x)′Σ−1 yy (y − VW ′x)〉 (2) = 1 2 tr(Iq)− tr(VW ΣxyΣ yy ) + 1 2 tr(WV ′Σ−1 yy VW Σxx) (3) = q 2 − tr(W ΣxyΣ yy V ) + 1 2 tr(V ′Σ−1 yy VW ΣxxW ) (4) Define U = Σ−1/2 yy V and S = Σ 1/2 xx W . Then, C = q 2 − tr(S ′Σ−1/2 xx ΣxyΣ yy U) + 1 2 tr(U ′US ′S) (5) Scalar Case For the m = 1 case, the arguments of the traces in (5) are all scalar. We examine separately the magnitudes and directions of the vectors U and S. Denote the magnitude of U by γu and the unit vector in the direction of U by eu, similarly for S. Then, C = q 2 − γuγsesΣ xx ΣxyΣ yy eu + 1 2 γ uγ 2 s . (6)

Read the paper · More papers on PaperTik