ICFHR2018 Competition on Automated Text Recognition on a READ Dataset

Tobias Strauß, Gundram Leifert, Roger Labahn, Tobias Hodel, Günter Mühlberger · 2018

We summarize the results of a competition on Automated Text Recognition targeting the effective adaptation of recognition engines to essentially new data. The task consists in achieving a minimum character error rate on a previously unknown text corpus from which only a few pages are available for adjusting an already pre-trained recognition engine. This issue addresses a frequent application scenario where only a small amount of task-specific training data is available, because producing this data usually requires much effort. We present the results of five submission. They show that the task is a challenging issue but for certain documents 16 pages of transcription are sufficient to adapt a pre-trained recognition system.

Read the paper · More papers on PaperTik