Arabic Video Text Recognition Based on Multi-dimensional Recurrent Neural Networks

Oussama Zayene, Soumaya Essefi Amamou, Najoua Essoukri BenAmara · 2017

In this paper we propose a novel method based on Recurrent Neural Networks (RNN) for text recognition in Arabic news video frames. In fact, embedded texts in videos represent a rich source of information for indexing and automatic processing of multimedia documents. However, text recognition in video is not trivial due to many challenges like background complexity (e.g., presence of text-like objects), unknown text size/font with various colours and degraded text quality. The proposed system presents a segmentation-free recognition technique (i.e. no prior required segmentation of words into characters) using a RNN architecture. This technique relies specifically on a Multi-Dimensional Long Short Term Memory (MDLSTM) with a Connectionist Temporal Classification (CTC) output layer. Our system has been evaluated on the public AcTiV-R dataset. The obtained results are very promising.

Read the paper · More papers on PaperTik