Recognition and connection of moving captions in Arabic TV news

Seiya Iwata, Wataru Ohyama, Tetsushi Wakabayashi, Fumitaka Kimura · 2017

The authors have conducted studies on Arabic news caption recognition to develop a system for video retrieval by keyword to index and edit Arabic broadcast programs received daily and stored in a big database. This paper proposes a dedicated OCR for recognizing low resolution news caption in video images. The news caption recognition system consisting of text line extraction, word segmentation and recognition of words is developed and the performance is experimentally evaluated using Dataset of frame images extracted from AlJazeera broadcasting programs. This paper also proposes a technique to connect the recognized moving news captions into a sentence. The proposed method is necessary for automatic language translation and is also capable of reducing the OCR errors due to truncated characters at both ends of the running news captions. The proposed connection method is a technique based on insertion operation with minimum edit distance between successive two news captions to be connected. Character likelihood based substitution method is newly proposed and comparatively tested with majority based substitution method. For the Dataset, character recognition rate (F-measure) after moving news caption connection by proposed method using bi-gram sequence (Method-B) realized simple processing and was improved to 98.74% from the rate 96.12% before connection processing.

Read the paper · More papers on PaperTik