On-line handwritten document understanding
Anil Kumar Jain, Anoop Namboodiri · 2004
This thesis develops a mathematical basis for understanding the structure of on-line handwritten documents and explores some of the related problems in detail. With the increase in popularity of portable computing devices, such as PDAs and handheld computers, non-keyboard based methods for data entry are receiving more attention in the research communities and commercial sector. The most promising options are pen-based and speech-based inputs. Pen-based input devices generate handwritten documents which have on-line or dynamic (temporal) information encoded in them. Digitizing devices like Smart-Boards and computing platforms, which use pen-based input such as the IBM Thinkpad TransNote and Tablet PCs, create on-line documents. As these devices become available and affordable, large volumes of digital handwritten data will be generated and the problem of archiving and retrieving such data becomes an important concern. This thesis addresses the problem of on-line document understanding, which is an essential component of an effective retrieval algorithm. This thesis describes a mathematical model for representation of on-line handwritten data based on the mechanics of the writing process. The representation is used for extracting various properties of the handwritten data, which forms the basis for understanding the higher level structure of a document. A principled approach to the problem of document segmentation is presented using a parsing technique based on stochastic context free grammar (SCFG). The parser generates a segmentation of the handwritten page, which is optimal with respect to a proposed compactness criterion. Additional properties of a region are identified based on the nature and layout of handwritten strokes in the region. An algorithm to identify ruled and unruled tables is presented. The thesis also proposes a set of features for identification of scripts within the text regions and develops an algorithm for script classification. The thesis presents a matching algorithm, which computes a similarity measure between the query and handwriting samples in a database. The query and database samples could be either words or sketches, which are automatically identified by the document segmentation algorithm. We also address the related problem of document security, where on-line handwritten signatures are used to watermark on-line and off-line documents. The use of a biometric trait, such as signature, for watermarking allows the recipient of a document to verify the identity of the author in addition to document integrity.