Text Detection and Style Classification from Images Using Vision Transformer and Transformer Decoder

Hideaki Yajima, Chee Siang Leow, Hiromitsu Nishizaki · 2024

This paper introduces a Transformer-based model for text detection and classification in images. By combining a Vision Transformer (ViT) encoder and a Transformer decoder, the proposed approach directly estimates bounding box coordinates of text regions, eliminating the need for post-processing. The ViT encoder captures local and global features, which the decoder uses to generate bounding boxes and classify text as handwritten or typeset text. Experiments on synthetic text images show the model achieves high accuracy in both detection and classification tasks, surpassing other methods. This technique could improve OCR accuracy and enhance the processing of documents containing both handwritten and typeset text.

Read the paper · More papers on PaperTik