Automated Document Processing: Combining OCR and Generative AI for Efficient Text Extraction and Summarization

G. Abinaya, Guttulapavan Durga Rao, Pragada Sri Ram Gopal, Manas Manglik Kaushik, B.P. Sreejith Vignesh · 2024

Technological advancement including the use of new media formats in creating digital documents calls for the development of enhanced processing techniques and methods. Old school Optical Character Recognition (OCR) tools remain useful but are inadequate when faced with sophisticated documents or low-quality scans. Also, summarizing the textual content of such papers continues to be a manual process and adds inefficiency and potential errors in information gathering. Finally, an integrated approach to the automation of both text extraction and summarization based on improved OCR and Google Gemini's generative AI technique is presented in this paper. By using more precise OCR algorithms, the system greatly increases accurate text recognition from poorly quality scans or complex structures. At the same time, the employment of generative AI models contributes to the formulation of brief and relevant summaries and also improves document retrieval and handling. Its effectiveness has thus been confirmed by testing on a variety of documents, and it has increased both accuracy and utility compared to previous techniques.

Read the paper · More papers on PaperTik