Videotext OCR using hidden Markov models
Prem Natarajan, Baback Elmieh, Richard M. Schwartz, J. Makhoul · 2002
We present a method for performing optical character recognition (OCR) of text in video images. Recognition of videotext is a challenging problem due to various factors such as the presence of rich, dynamic backgrounds, low resolution, color, etc. Our strategy is to process the video images to produce high-resolution binarized text images that resemble printed text. We describe a novel clustering and relaxation procedure that combines stroke and color information to separate the text from the background. The binarized text image is then recognized with our Byblos OCR engine (Natarajan et al., 2001; Schwartz et al., 1996) using hidden Markov models trained on similar data. We present experimental results on a video-data corpus collected from broadcast news programs. Currently the system delivers a character error rate of 8.3% on independent multi-font test data from this corpus.