Assessing Generalization by 2-D Receptive Field Visualization
Michael K. Arras, Peter Protzel · 1993
A performance comparison of different neural networks for classification tasks has to take the convergence speed on the training data as well as the generalization capability on unseen test data into account. We present visualization results for two two-dimensional benchmark problems by scanning the receptive fields of different networks. This allows for the assessment of generalization and shows the evolution of the networks during training. The figure below shows the final receptive fields for three different networks that learn to distinguish between two classes arranged as an 8x8 checkerboard a), b) and c) and two intertwining spirals d), e) and f). The training points of the two classes are indicated by a cross and a square. The networks evaluated are a) a 2-12-10-1 MLP trained with backpropagation, d) a 2-5-5-5-1 MLP with fully connected “cross-cut” connections trained with backpropagation, b) and e) a standard Cascade-Correlation network and c) and f) a Cascade-Correlation network that uses a mixture of sine, cosine and sigmoid activation functions. Note that the number of epochs it took to solve the 8x8 checkerboard problem with a similar generalization differs by almost two orders of magnitude between a) and c). In contrast to c), while the use of mixed activation functions offers a similar speed advantage, the resulting generalization in f) is very poor. Our results show that convergence speed, network size and generalization performance do not necessarily go hand in hand, and that visualization is a valuable aid which should be part of any benchmark that assesses the performance of different network architectures. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.