Automatic Zoning of Digitized Documents
Heath Nielson, William A. Barrett · 2001
Recent improvements in scanning technology have made available (over the web) millions of scanned genealogical documents. However, in order to exploit the content of these documents, the granularity of the indexing must move from the image level to individual fields within the document. Being able to search or browse through individual fields of a document rather than the whole image is the first step in indexing and understanding the content of those fields (e.g. name, age, sex, location, etc.). In addition, field-level addressing gives us a means of partitioning the document into meaningful, and relevant components with all of the other attendant benefits of speed and the economy of data transfer and storage. Rather than transferring and searching through the entire document, selected fields could be transmitted instead. This allows users to focus only on the information important to them. Sending portions of an image also provides for a quicker response time allowing the user to efficiently locate the information they are searching for. Segmentation of a document into its respective fields also allows each field's contents to be contextually analyzed. For example, a field that contains printed text would be sent to an OCR engine. Fields containing handwriting would be stored for subsequent semi-automated or user-assisted interpretation. To perform automated field-level indexing and addressability, automated zoning techniques are needed to automatically partition the document and identify the location and content of regions and fields. We have developed a zoning algorithm which allows rectangular regions of interest in a document to be identified and partitioned. We also propose methods to determine whether these regions contain handwriting or printed text.