Classification Capabilities of Architect we-specific Recurrent Networks

Jacques Ludik, Ian Cloete · 1995

The classification capabilities of Elman and Jordan architecture-specific recurrent threshold networks are analyzed in terms of the number and possible types of cells the networks are capable of forming in the input and hidden activation spaces. For Elman networks the number of cells is always 2h, there are no closed or imaginary cells, and they are therefore not capable of forming disconnected decision regions. For Jordan networks this is only the case when the number of hidden units are less or equal to the sum of input and state units. We have interpreted the equations obtained, compared the results with feedforward threshold networks, and illustrated them with an example. In this paper the classification capabilities of the two most popular architecture-specific recurrent neural networks (simple recurrent networks), namely Elman [l] and Jordan networks [2], are analyzed. We examine the number and possible types of cells in the input and hidden activation space for classification as a function of the number of network units. The basic functioning of a multilayer neural network is that the first hidden layer uses hyperplanes to partition the input space into a number of cells. The main function of the additional layers, in particular the output layer, is then to group these cells into larger regions. A cell [4] is a polyhedric region in the input activation space, which are labeled by the hidden activation values that specify on which side of each hyperplane the points in that region lie. A cell in the input space has a corresponding cell in the hidden space, both which are labeled by the same hidden activation values. An open cell in the input space has input variables which are unbounded, whereas a closed cell in the input space has input variables which are always bounded. An imaginary cell is labeled by hidden activation values that have no corresponding input values and appears as a cell in the hidden space, but not in the input space. To distinguish between cells that are formed in the input space and those that are imaginary, we introduce the term real cells to denote the former. The number of real cells is the sum of the open and closed cells. When the term cell is used on its own, it indicates a real cell. For the theoretical analysis, we consider networks where the unit nonlinearities are threshold elements with binary outputs, i.e. units with a step threshold activation function. We denote a feedforward threshold network (FFTN) with n input units, a hidden layer of h threshold units, and s output threshold units by n:h:s FFTN. An Elman architecture-specific recurrent threshold network (ASRTN) with n input units, h hidden threshold units, a context layer of h linear units, each containing a copy of a corresponding hidden threshold unit value on the previous time step, and s output threshold units, is denoted by n:h:s Elman ASRTN. An n:h:s Elman ASRTN can be viewed as an (n+h):h:s FFTN having h special additional inputs, the context units. A Jordan ASRTN with n input units, h hidden threshold units, s output threshold units, and a state layer of s linear units, which contain a linearly averaged copy of previous output values (a proportion of its previous values plus the previous output values), is denoted by n:h:s Jordan ASRTN. An n:h:s Jordan ASRTN can be viewed as an (n+s):h:s FFTN having s special additional inputs, the state units. The conclusions of this analysis can be extended to networks with the well-known sigmoid activation function. This can be accomplished by a sigmoid function that approximates a threshold function or just by multiplying all the weights and thresholds in the network by a sufficiently large constant, forcing the outputs to be close to one or zero (the gain parameter is large). In the following sections we discuss the number and types of cells for Elman and Jordan ASRTNs, compare the results with FFTNs, illustrate them with an example, and interpret the equations obtained.

Read the paper · More papers on PaperTik