Robust, Simple Page Segmentation Using Hybrid Convolutional MDLSTM Networks

Thomas Michael Breuel · 2017

Analyzing and segmenting scanned documents is an important step in optical character recognition. The problem is difficult because of the complexity of 2D layouts, the small tolerance of segmentation errors in the output, and the relatively small amount of labeled training data available. Traditional approaches have relied on a combination of sophisticated geometric algorithms, domain knowledge, heuristics, and carefully tuned parameters. This paper describes the use of deep neural networks, in particular a combination of convolutional and multidimensional LSTM networks, for document image and demonstrates that relatively simple networks are capable of fast, reliable text line segmentation and document layout analysis even on complex and noisy inputs, without manual parameter tuning or heuristics. The method is easily adaptable to new datasets by retraining and an open source implementation is available.

Read the paper · More papers on PaperTik