Document Image Page Segmentation and Character Recognition as Semantic Segmentation
Seth Stewart, Bill Barrett · 2017
Convolutional Neural Networks (CNNs) have produced excellent results in natural scene semantic pixel labeling tasks. We examine the application of this idea to document processing, using fully supervised Deep CNN semantic segmentation to separate content layers from historical document images containing diverse content types, including handwriting, machine print, form lines, and stamps. For efficiency, we employ a downsampling-upsampling network to make dense pixel predictions. CNNs achieve high generalization accuracy on document images with interleaved, overlapping strokes, even when trained on a solitary pixel-labeled document image. We also show a proof-of-concept extension of the semantic segmentation task to handwritten cursive character recognition, enabling a new "segmentation-free" approach to handwriting transcription.