Enhancing OCR Accuracy for Bugis Language to Bahasa Indonesia Dictionary Conversion through Image Pre-processing and Scaling Techniques

Mohammad Teduh Uliniansyah, Agung Santosa, Rachmawan Atmaji Perdana · 2024

Optical Character Recognition (OCR) systems often struggle with low-resource languages like Bugis due to script complexity and limited digitized resources. This research aims to improve OCR accuracy for Bugis–Indonesian dictionary conversion by applying advanced image pre-processing techniques. We evaluated the performance of Tesseract-Vanilla, Tesseract with ImageMagick, Python PIL, and ChatGPT's OCR models by measuring Character Error Rate (CER) and Word Error Rate (WER). The results show that Tesseract with ImageMagick significantly outperformed other methods, reducing CER from 8.94% to 2.82% and WER from 35.35% to 13.54%. These findings highlight the effectiveness of pre-processing techniques in enhancing OCR performance for low-resource languages like Bugis.

Read the paper · More papers on PaperTik