Combining deep learning and language modeling for segmentation-free OCR from raw pixels

Stephen Rawls, Huaigu Cao, Ekraam Sabir, Prem Natarajan · 2017

We present a simple yet effective LSTM-based approach for recognizing machine-print text from raw pixels. We use a fully-connected feed-forward neural network for feature extraction over a sliding window, the output of which is directly fed into a stacked bi-directional LSTM. We train the network using the CTC objective function and use a WFST language model during recognition. Experimental results show that this simple system outperforms extensively tuned state-of-the-art HMM models on the DARPA Arabic Machine Print corpus.

Read the paper · More papers on PaperTik