Printed Cyrillic character recognition system
J.E. Tierney, N. Revell · 1994
This paper describes an off-line system for recognising handprinted or machine printed non-cursive Russian characters-the Cyrillic alphabet, including basic numerals and punctuation used in the Russian language. The input is the scanned text stored in a data file (using TIFF format), and the output is a stream of codes representing the characters identified as a text file. The system comprises a graphics file reader, segmenter, feature extractor, classifier and checker. The feature extractor generates a feature pattern from a combination of structural information describing key character strokes and statistical information describing the distribution of points in the character. The classifier is a multi-layer feedforward neural network which assigns classification to each feature pattern after training under supervised learning using a backpropagation learning rule. The checker maps the classifications to character codes, and provides the framework for additional language specific rules. The system is independent of any particular operating system, programming language compiler, or hardware platform.>