Self-Attention Enhanced Recognition: A Unified Model for Handwriting and Scene-text Recognition with Improved Inference

Gaurav Patel, Taewook Kim, Qian Lin, Jan P. Allebach, Qiang Qiu · Electronic Imaging · 2024

In this paper, we introduce a unified handwriting and scene-text recognition model tailored to discern both printed and hand-written text images. Our primary contribution is the incorporation of the self-attention mechanism, a salient feature of the transformer architecture. This incorporation leads to two significant advantages: 1) A substantial improvement in the recognition accuracy for both scene-text and handwritten text, and 2) A notable decrease in inference time, addressing a prevalent challenge faced by modern recognizers that utilize sequence-based decoding with attention.

Read the paper · More papers on PaperTik