Robustness analysis of neural networks with an application to a neuro-controller problem

Kalmanje S. Krishnakumar, K. Nichita · Guidance, Navigation and Control Conference · 1996

This paper analyzes the robustness characteristics of multi-layered feed-forward Neural Networks (NN) using linear systems theory. Robustness of NN is defined as the NN's ability to perform within certain bounds of its nominal (without uncertainty) performance in the presence of bounded uncertainty. An induced Euclidean matrix norm is used to derive error bounds for NN with activation functions that predominantly exhibit linear behavior. Lyapunov stability theory is used to derive bounds on the non-linear variations in NN with activation functions that exhibit non-linear behavior. A Monte Carlo simulation analysis is conducted to examine the robustness characteristics of fully and sparsely connected networks. The following conclusions are drawn based on the above analysis: (a) sparsity in the NN connection topology is highly desirable to achieve robustness; (b) two hidden layer networks with equal number of neurons in each layer exhibit very poor robustness; (c) a fully forward-connected network with sparsity is the most robust and accurate for a given number of neurons; and (d) for NN with many neurons, activation functions with highly non-linear regions exhibit poor robustness. A system identification problem and a neuro-controller application to the pitch attitude control of a space station model are presented to reinforce the results of the robustness analysis. Introduction Robustness of neural networks (NN) is an important characteristic for many applications ranging from simple function mapping to more complicated non-linear control problems. Robustness of any system can be defined as the system's ability to perform within certain bounds of its nominal (without uncertainty) performance in the presence of bounded uncertainty. In the NN context, robustness relates to bounded output error of the NN to bounded input errors. These inputs are seen as inputs to any neurons and outputs are seen as outputs from any neurons. By considering inputs and outputs from individual neurons instead of the whole NN, we can include uncertainties related to structural failures (such as loss of nodes, weights, etc.), internal noise (hardware-related) in the NN, and also examine uncertainties in the inputs to the NN system. Where does generalization fit into all of these? Generalization of a NN is its ability to sensibly interpolate input patterns that are new to the network [1]. Generalization is a loosely used terminology. Training a NN for a desired generalization is impossible due to the simple reason that we do not know the function to be mapped a priori. To illustrate this, we present in Figure 1 a set of training data pairs [x, y]. It is obvious to see that one can have infinite number of lines go through these points. Which one is the best? Since this is unknown, how can one predict a priori the generalization characteristics of the NN? This can only be done approximately by using test data after the training is done. In conclusion, generalization depends to a great extent on the training set used and is also difficult to define a priori. Whereas, robustness is easy to define using input-output uncertainty bounds as shown in Figure 1. Most popular neural networks in use today use mullet-layered feed-forward networks with connections going from one layer to the next. Several investigators have shown the approximating capabilities of these networks. Another type of network that has received some attention is the fullyforward connected network [2,3]. KrishnaKumar [3] has shown using empirical results that if this type of network is made sparse, the networks are more accurate and robust for a given number of neurons. Many other researchers have shown similar empirical results using sparse multi-layered feedforward networks (Mozer et al. [4J, Rumelhart et al. [5], and Sietsma et al [6]). Copyright © 1996 by K. KrishnaKumar. Published by the American Institute of Aeronautics and Astronautics, Inc. with Permission. In this paper, we address the general question of how to define relative robustness of different neural network structures. We assume that the network has been trained to provide a desired accuracy for the training data set. The question then is simply one of how to relate the given NN structure to certain bounds on its mapping error for untrained data. We first define robustness given a NN structure and then examine the robustness of layered networks, fully-forward connected networks, and sparse networks. The paper is organized as follows. We begin with the definitions of NN structures that are discussed in this paper. Next, we state the linear systems theory results that will be used to define robustness. We then present an interpretation of NN equations in terms of the linear systems theory and present a Monte Carlo simulation analysis to examine the robustness characteristics. The following conclusions are drawn based on the above analysis(a) sparsity in the NN connection topology is highly desirable to achieve robustness; (b) two hidden layer networks with equal number of neurons in each layer exhibit very poor robustness; (c) a fully forward-connected network with sparsity is the most robust and accurate for a given number of neurons; and (d) for NN with many neurons, activation functions with highly non-linear regions exhibit poor robustness. Finally, a system identification problem and a neurocontroller application to the pitch attitude control of a space station model are presented to reinforce the results of the robustness analysis. Neural Network Preliminaries In this paper, without loss of generality, we will assume one input, one output neural networks with N-2 sigmoidal hidden neurons. The hidden neurons could be arranged either in a layered form or in a fully-forward connected form (Figure 2). A fully-forward connected NN could be seen as a layered NN as shown in Figure 3. Also, the fullyforward connected NN with the proper connections removed is equivalent to a traditionally layered form. The equation for a fully-forward connected NNs given as xi = NN input and the input neuron is a linear neuron. In the above equations, y; = output of the i* neuron Wjj = weight connecting the i* neuron to the j* neuron. i ( ) = non-linear activation function. Equation 1 can be used to arrive at layered networks by zeroing out appropriate elements of the weight matrix. This will be illustrated later. Mathematical Preliminaries Vector Norm: For any vector x e R, the Frobenius (Euclidean) norm is given as ____

Read the paper · More papers on PaperTik