Statistical risk analysis for classification and feature extraction by multilayer perceptrons

Chanchal Chatterjee, Vwani Roychowdhury · 2002

We investigate the training of multilayer perceptrons with the commonly used mean square error (MSE) criterion, and demonstrate a number of novel connections between the neural network operations and the Bayes risk analysis, Although previous research shows a number of connections from seemingly different criteria, we establish a common statistical framework to derive a generalized version of most, if not all, of these results, and also present several new results. We discuss the following: (1) We present two equivalent cost functions, and show that the MSE at the network output is equivalent to these cost functions for large samples. (2) We show that if the network performs a weighted classification, then the network output estimates the conditional risk. (3) We next show that if the final layer of the network is linear, then minimizing the MSE at the output, also maximizes a generalized criterion for nonlinear discriminant analysis (NDA). (4) We show that for a network with linear output layer, the outputs sum to one, and behave like probabilities. This new result allows us to estimate conditional risks at the network output, and also perform NDA at the final hidden layer. (5) Results for the uniform costs show that the MSE at the output is a tight upper bound of the error probability of the Bayes decision rule.

Read the paper · More papers on PaperTik