TSEP: text spotting based on transformer with bidirectional explicit points sampling

Mingqiang Luo, Jin Huang, Jianbo Li, Haijie Wang, Yu Chen, Liuran Ren, Wanyu Ji · 2025

Recent years, the Transformer has achieved remarkable results in scene text spotting. This paper uses the classic encoderdecoder architecture to propose a Text Spotting model based on Transformer with bidirectional Explicit Points sampling (TSEP). As a sequence, the text contains rich semantic information in its forward and backward features. We model each text instance through bidirectional explicit point sampling. After decoding by the decoder, the positional and semantic information of the text is integrated into the explicit points. Therefore, a basic prediction head is capable of producing the boundary of the text region, the text content, and the corresponding confidence scores. Additionally, we propose a reference point feature enhancement module constructed using one-dimensional convolutions and MLP to address the spatial inductive bias of non-local self-attention in Transformers. Experimental results across various public datasets indicate that our model outperforms numerous other leading models.

Read the paper · More papers on PaperTik