PIFText: Progressive Injection Fusion for Text Detection in Autonomous Driving Scenarios

Chenyu Yuan, Jing Zhang, Jiafeng Li, Zhuo Li · 2025

In the contemporary field of intelligent transportation, autonomous driving (AD) task exhibits a significant demand for environmental comprehension, and scene text is a key element in improving environmental comprehension. The inherent complexity of the surrounding environments under AD gives rise to a diversity of scene text geometries, thus making the text detection task particularly challenging. To address this issue, we propose a novel progressive injection fusion for text detection (PIFText) in AD scenarios. Initially, we propose a progressive injection-feature pyramid network (PIFPN), a novel architectural module specifically designed to mitigate information loss or quality degradation during multi-level feature transmission while effectively addressing the substantial semantic gap between non-adjacent hierarchical levels. Subsequently, we design an enhanced feature feed-forward network (EFFN) architecture to efficiently capture local information through depthwise convolution (DWConv) and partial convolution (PConv) operations. In addition, we introduce a multi-head deformable attention mechanism that enables the model to dynamically focus on relevant regions while accelerating model convergence. We conduct extensive experiments on publicly available ICDAR19-ArT, Total-Text and CTW1500 datasets. The results shows that our PIFText achieves precisions of 87.0%, 93.1%, and 91.9%, recalls of 73.2%, 87.5%, and 87.3%, and F-measures of 79.5%, 90.1%, and 89.4% on the three datasets, respectively, demonstrating its effectiveness and superiority.

Read the paper · More papers on PaperTik