Rotation and scale invariant pattern recognition using a multistaged neural network

Jay I. Minnix, Eugene S. Mcvey, Rafael M. Inigo · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 1991

This paper presents a pattern recognition system that self-organizes to recognize objects by shape. The images are processed using a log-polar transformation that maps rotations and magnifications into representative translations. The systems then uses a multistaged hierarchical neural network that exhibits insensitivity to translations in representation space, which corresponds to rotations and scalings in the image space. The network's three layers perform the functionally disjoint tasks of preprocessing (dynamic thresholding), invariance (position normalization), and recognition (identification of the shape). The Preprocessing stage uses a single layer of elements to dynamically threshold the grey level input image into a binary image. The Invariance stage is a multilayered neural network implementation of a modified Walsh-Hadamard transform that generates a representation of the object that is invariant with respect to the object's position, which maps back to an invariance to rotational orientation and/or size. The Recognition stage is a modified version of Fukushima's Neocognitron that identifies the normalized representation by shape. The resulting network can successfully recognize objects that have been rotated, scaled, or a combination of both. The network uses a small number of fairly simple elements, a subset of which self-organize to produce the recognition performance.

Read the paper · More papers on PaperTik