Opportunities and Challenges of LLMs as Post-OCR Correctors

Radoslav Koynov, Triet Ho Anh Doan · Annals of Computer Science and Information Systems · 2025

Large Language Models (LLMs) have demonstrated potential as zero-shot Post-OCR correctors for historical texts.However, previous research has typically focused on a single data set and only evaluated Character Error Rate (CER) or Word Error Rate (WER).This study investigates the potential of LLMs to enhance the accuracy of Optical Character Recognition (OCR) and the limitations of the models.To this end, an evaluation of the approach is conducted for a number of German and English historical datasets, with an in-depth analysis of the model corrections and deviation from the ground truth.We demonstrate that LLMs have the capacity to enhance the quality of OCR results as zero-shot correctors in some cases, and finetuning LLMs shows promise as part of an LLM-based Post-OCR correction system, if certain risks are mitigated carefully.

Read the paper · More papers on PaperTik