Extraction of Text from Images using Document Understanding Transformer (Donut)

Vijaya Babu Panthagani, Mandadi Srinivas Reddy, Bonam Praveen Kumar, Shaik Irshad · 2024

Comprehending visual document images, like bills, is a challenging task that necessitates text extraction and a thorough comprehension of the document’s contents. This is addressed by visual document understanding (VDU), which extracts text using readily available optical character recognition (OCR) engines and concentrates on interpreting the output. Although OCR-based methods have shown encouraging results, they have drawbacks such expensive processing, rigid models, and the possibility of errors extending to later phases. This paper presents the Document Understanding Transformer (Donut), a novel OCR-free VDU model. This study describes a basic Transformer architecture with a pre-training aim of cross-entropy loss as part of the initial phase of OCR-free VDU research. Donut is a simple and efficient solution. Numerous tests and analyses show that Donut, a straightforward OCR-free VDU model, works faster and more accurately than other models for particular VDU tasks. To facilitate versatile model pre-training across several languages and areas, we offer a generator of synthetic data.

Read the paper · More papers on PaperTik