iDocChip
Vladimir Rybalkin, Syed Saqib Bukhari, Muhammad Mohsin Ghaffar, Aqib Ghafoor, Norbert Wehn, Andreas R. Dengel · 2018
End-to-end Optical Character Recognition (OCR) systems are heavily used to convert document images into machine-readable text. Commercial and open-source OCR systems (like Abbyy, OCRopus, Tesseract etc.) have traditionally been optimized for contemporary documents like books, letters, memos, and other end-user documents. However, these systems are difficult to use equally well for digitizing historical document images, which contain degradations like non-uniform shading, bleed-through, and irregular layout; such degradations usually do not exist in contemporary document images.