CNN for historic handwritten document search

Lech Szymanski, Steven J. Mills · 2017

While character recognition for printed text and handwriting in restricted cases is well studied, collections of historic documents provide a number of additional challenges. These include the complexities of cursive script, limited training data, and degradation of the documents themselves. We present initial work in whole-word recognition on the Marsden document archive held by the Hocken Collections in Dunedin, New Zealand. Based on transcriptions, but no word or character localisation information, we show that CNNs are a promising avenue for document search in such collections.

Read the paper · More papers on PaperTik