Enhancing Medical Records Digitization Through a Post-OCR Processing Technique
Mohamed Mehfoud Bouh, Forhad Hossain, Prajat Paul, Ashir Ahmed · 2024
In an era marked by technological progress, many healthcare service providers in developing countries still rely on paper-based records for storing patients' medical information. About 80% of medical data remains unstructured, including paper and digitized records that are hard to interpret. Optical Character Recognition (OCR) technology is essential in digitizing analog medical records. However, post-OCR accuracy is significantly influenced by the document layout and the order of detected bounding boxes. This paper addresses the digitization of medical records through OCR, focusing on reorganizing OCR-detected bounding boxes to ensure text is extracted accurately in a readable format. The method was tested on haematology and kidney reports using both open-source (Tesseract) and commercial (Google Vision) OCR tools. Results were compared to available multimodal models, specifically ChatGPT-4o, Claude3.5-Sonnet, and Gemini Advanced. The results demonstrate significant efficiency, with Character Error Rate (CER) of 0.05 and 0.12, and Word Error Rate (WER) of 0.07 and 0.21 for simple and complex documents, respectively.