Bilingual Road Text Recognition Based on a Hybrid Model of CTC and Attention

Ridha Zaghdoud, Khalil Boukthir, Tarek M. Hamdani, Adel M. Alimi · 2024

The recognition of Arabic and Latin text for autonomous cars involves developing systems and algorithms capable of accurately detecting and understanding Arabic characters and words from images. This technology is crucial in enabling autonomous vehicles to interpret and respond to Arabic and Latin traffic signs, road markings, and other textual information on the roads. The development of a reliable identification system, particularly for Arabic is challenging if a dataset contains differences in text size, typefaces, colors, orientation, illumination and noise. These problems become more difficult to solve. By exploiting the benefits of both CTC (Connectionist Temporal Classification) and Attention mechanisms to improve the accuracy and robustness of the recognition system, a hybrid CTC and Attention model is used in the prediction stage for Arabic-Latin image text recognition. Current research is focusing intensively on text panels in Latin, while other scripts, such as Arabic, remain little valued. For this reason, the NaSTSArLaTR (Natural Scene Traffic Sign and Panel Guide Arabic-Latin Text Recognition) dataset has been set up to validate our experiments. Our tests found that the suggested hybrid CTC and Attention model outperformed either Attention or CTC alone in the prediction phase. The dataset is publicly available in IEEE DataPort https://dx.doi.org/10.21227/phyg-_jc98.

Read the paper · More papers on PaperTik