Teaching computers to read handprinted paragraphs
Michael D. Garris · 1996
A set of algorithms is presented that enables a computer to automatically read a paragraph of handprinted text.These steps are broken down into: 1.) isolating the lines of handprint; 2.) segmenting the lines into character images; 3.) classifying the character images; and 4.) spell-correcting the classifications.A robust method of line reconstruction is presented that avoids statistical outliers and generates a spatial map representing the center of mass of each handprinted line.Components are organized into lines based on their overlap and proximity to bands within the map.This facilitates the efficient detection and splitting of vertically touching characters across lines of text A new character segmentor is also presented that adapts automatically to writing style.The segmentor composes characters from mul- tiple components and it separates touching characters from single components.Comparing the performance of the seg- mentor between a 1 -line response and a paragraph demonstrates that a person's writing has a tendency to become more varied and irregular as spatial constraints are relaxed.An optimized Probabilistic Neural Network is used to classify the character segments, and a description of a dictionary-based correction scheme is provided.Recognition accuracies are reported on 500 handprinted paragraphs from NIST Special Database 1 9 containing the Preamble to the U.S. Constitution.The new methods of line reconstruction and character segmentaticai improve word recognition accuracy by 4% (from 60% to 64%) over a system that uses connected components directly as character segments.This paper rep- resents a significant step toward a general purpose rmconstrained handprint recognizer, and points out where future work may be focussed.