Interpreting handwritten text in a constrained domain
Edward A. K. Cohen · 1992
This thesis explores the automatic interpretation of handwritten text when constraints exist on its content and structure. Interpretation of text is regarded as the derivation of textual content in the form of a predefined symbolic representation. The handwritten text is assumed to be off-line, in that the computer input is assumed to be from a scanned and digitized image of handwriting on paper, rather than input from a specialized on-line device such as a bit-pad. Examples of such handwritten text are addresses on envelopes, amounts on bank checks, and drug and dosage on drug prescriptions. Both the content and structure of the text are assumed to be constrained. Content constraints are assumed to exist in the form of: (i) lexicons that can be associated with syntactic categories (e.g., state names, account numbers, personal names), and (ii) known relationships (e.g., semantics, world knowledge) that exist between phrases in different syntactic categories. Phrases that correspond to syntactic categories are assumed to follow structural constraints. Structural constraints describe the text's two-dimensional phrase layout (e.g., the position of a phrase in a text line, the position of a text line in a text block). Writing style is assumed to be unconstrained, in that it consists of what is normally encountered in practice without placing additional restrictions on how individual characters are formed. A solution for this interpretation problem is described by a computational theory that consists of five stages: creating phrase hypotheses, computing visual features, categorizing phrases with high level processing, additional discriminating of phrase categories, and extracting the interpretation. The computational theory uses these five stages to describe an algorithm in which early recognition of primitives guides the location of phrases. Additional context is then used to identify relevant phrases and to derive an interpretation. The effectiveness of the theory is demonstrated in two application domains: handwritten postal addresses and handwritten bank checks. In each application, techniques described by the theory allow the use of additional context to improve performance, thereby justifying the theory.