DITOCR: A Decoder-Only Transformer for Industrial Optical Character Recognition

Shengyuan Wang, Wenjie Cai, Rihui Xia, Xingbo Dong, Zhe Jin · 2025

Industrial Optical Character Recognition (OCR) poses significant challenges due to factors such as low contrast, background clutter, uneven illumination, and distorted or occluded characters. In this paper, we propose DITOCR (Decoder for Industrial Character OCR Recognition), a novel decoder-only framework designed specifically for industrial text recognition tasks. Unlike traditional encoder-decoder architectures, DITOCR eliminates the need for a vision encoder by utilizing a lightweight convolutional frontend, a ViT-style patch embedding module, and a GPT-2-based Transformer decoder to perform end-to-end autoregressive text generation. To enhance robustness under realworld industrial conditions, we introduce a character-level attention mechanism that adaptively focuses on relevant spatial regions within the visual input. The model is trained and finetuned on a combination of synthetic and real industrial datasets, and achieves superior accuracy and character error rate (CER) compared to existing methods. Extensive experiments demonstrate that DITOCR offers a compact, effective, and accurate solution for industrial OCR applications.

Read the paper · More papers on PaperTik