Comparative Analysis of Large Language Models for OCR Post-Processing in Persian: From ParsBERT to GPT
Fatemeh Valizadeh, Fahimeh Ghasemian, Elham Shabaninia · 2025
Optical Character Recognition (OCR) refers to the automatic identification of text in images and its conversion into searchable and editable formats. Due to its extensive applications, OCR is considered a crucial and challenging topic in the field of computer vision. In Persian, the unique characteristics of the script often result in OCR outputs with significant errors, which can compromise readability and comprehension of the content, emphasizing the need for error correction, particularly in terms of spelling. This study investigates the performance of four large language models-ParsBERT, LLaMA, Mistral, and GPT-for enhancing OCR outputs. Additionally, an innovative approach that integrates ParsBERT with other models is introduced. These models were evaluated on three corpora of varying sizes and complexities using diverse evaluation metrics, including precision, accuracy, recall, and others. Our comparative analysis highlights the strengths and weaknesses of each model across different scenarios, providing deeper insights into their performance. Furthermore, the results demonstrate that the proposed hybrid approaches outperform standalone models across all corpora, significantly reducing error rates and improving key metrics such as precision, accuracy, and recall.