Optimization of OCR in Detecting Research Proposal and Lecturer Community Service Documents using Thresholding Method

Aghus Sofwan, Annisa Yasmin Sumardi, Imam Adhi Santoso, Yosua Alvin Adi Soetrisno, Muhammad Arfan, Eko Handoyo · 2021

The development of science and technology today has changed various aspects of life, such as digitizing documents. Document digitization makes it easier to archive and search for documents ranging from book documents, ancient manuscripts, and administration. One method to extract image text into text that a computer can process is optical character recognition (OCR). However, OCR has a problem when the scanned file has noise such as blur, varying light intensity, and tilt also OCR cannot process files with the pdf extension. Therefore, image quality improvement is carried out by optimizing the thresholding method in the preprocessing process. The OCR development process uses the python programming language and the wand library to convert documents with pdf extensions into images. Optimization of thresholding in preprocessing is done by testing nine methods with the best results, namely using trunc thresholding with an accuracy of 89%. Then the pytesseract library to extracts texts in the image into editable text. The feature creation using the OCR method has been through accuracy testing using the Levenshtein distance algorithm with 95.3 % results. The performance testing with an average time of 15.2 s to produce a product ready to be integrated with the Research and Lecturer Community Service Information System (SITEDI), Faculty of Engineering, Diponegoro University and is expected to facilitate lecturers in filling out the proposal submission form.

Read the paper · More papers on PaperTik