Enhancing OCR Model Training with Synthetic Data: A Case Study with SROIE Dataset

Yuriy Yashchuk, Yuliana Yurchenko, Nataliia Yurchenko, Yurii Andrashko · 2024

This paper investigates the impact of synthetic data on the training effectiveness of optical character recognition (OCR) models. With the rapid advancement in machine learning, the availability of diverse and extensive datasets has become crucial for developing robust OCR systems. However, the scarcity of annotated real-world data poses significant challenges. This study focuses on performing an OCR task using the SROIE dataset, which presents a real-world challenge due to prevalent issues such as paper and ink corruption, inconsistent lighting conditions, variable fonts, colors, background noise, and imaging distortions. We introduce an approach to generating synthetic data, aiming to mimic those peculiarities. By conducting thorough experiments using the SROIE dataset and a custom synthetic training dataset, we demonstrate substantial improvements in model performance, thereby underscoring the potential of synthetic data in OCR applications.

Read the paper · More papers on PaperTik