Using ChatGPT as a Combined Invoice OCR and Key-Value Extractor
Nemi Pelgrom, Morgan Ericsson, Johan Hagelbäck, Jonas Nordqvist, Håkan Grahn · 2025
This paper provides details and analysis of findings from three experiments on the capabilities of the OpenAI chatbot ChatGPT-4 with vision capabilities as a dual-purpose tool for Optical Character Recognition (OCR) and key-value data extraction tasks. Our results could be relevant to any task where one is interested in extracting key information from images with approximately equal complexity, and are scalable to be used for large datasets. We did the main experiments in the OpenAI user interface for the model, and a smaller experiment using the API. The experiment used a dataset comprising 1000 digital invoices alongside 1000 photographic images of real receipts, collected from a broad spectrum of market sectors within Sweden. The main experiment gave us a significant accuracy rate, achieving$\mathbf{9 9. 8 \%}$in extracting critical financial information from the digital invoices. Similarly, when applied to the photographic images of receipts, it maintained a high accuracy level of$\mathbf{9 9. 5 \%}$. The smaller experiment gave us an accuracy of$\mathbf{9 4. 4 \%}$These findings are particularly noteworthy not only because of the high accuracy levels but also due to the model's effectiveness in performing both OCR and key-value extraction tasks as a one-step process. This dual functionality underscores the model's potential as a highly efficient and reliable solution in automating financial data extraction and processing tasks.