Algorithms for extracting text from degraded document images
Yan Chen · 2007
Documents usually contain a large amount of information and have been the primary information medium in our society.From handwritten or printed letters and words, signed cheques and contracts in our normal life to the covenants between countries, documents are an important medium of record and their importance in law still cannot be replaced by any other medium.For newly created electronic documents, searching based on keywords or phrases is relatively straightforward as the documents are created using appropriate software which makes them easily compatible with other software enabling keyword searching to be readily performed.However, many older documents only exist in paper form and are usually converted to computer form by scanning the documents and storing them as images in appropriate formats.Popular image formats are Adobe pdf, Postscript and TIFF.Images scanned in these formats can only be displayed or printed using computer tools.It is not possible to search them by keywords unless sophisticated image-processing (such as image segmentation, image layout analysis, image understanding and image classification) tools are applied.This makes document image analysis an important research area.The main objective of this research is to automatically separate text from the background in degraded scanned document images and locate individual words in the text.In this thesis, techniques are presented to extract the handwritten text from the noisy or degraded background.This is accomplished through a multi-stage technique, which analyses the feature vectors in a local block and then chooses the most appropriate threshold method in a database for each block.The multi-stage algorithm is suitable for ii