An Enhanced Vision-Language Pre-Training Approach for Scene Text Detection
Tao Gu, Chongyang Zhang · 2024
Recent advancements in Vision-Language Pre-training (VLP) techniques have greatly improved performance in Scene Text Detection tasks by leveraging the rich visual and textual content in scene text images. We propose an innovative integration of contrastive learning with Masked Language Modeling (MLM) and Masked Image Modeling (MIM), inspired by the Masked Autoencoder (MAE) approach. This synthesis enhances self-supervised learning by combining discriminative and generative learning. Notably, we introduce masked image modeling for text detection, enabling effective representation learning in text image regions. To distinguish our approach from traditional BERT's MLM, we develop a custom tokenizer tailored for scene text detection. Our pre-trained model aligns visual and textual information, resulting in improved performance of existing scene text detectors. Extensive experiments on datasets such as ICDAR2015 show that our method significantly outperforms previous pre-training approaches.