Structure-preserving document image compression

Omid E. Kia, David Doermann · 2002

Maintaining a document in image form is often preferable in order to avoid the high cost of manual conversion or the introduction of large numbers of errors by automatic OCR and/or graphics interpretation. The large volume of data in the image can be greatly reduced by using compression techniques. Text-intensive document images typically have a great deal of redundancy in the bitmap representations of symbols, and we make use of that redundancy for compression by clustering components, representing each cluster by a template and encoding the error. Our method is novel in modeling the error associated with each cluster and in preserving structure, an important component for readability and processing.

Read the paper · More papers on PaperTik