Using OCR Framework and Information Extraction for Thai Documents Digitization
Todsanai Chumwatana, Waramporn Rattana-umnuaychai · 2021
At present, digital transformation has taken into an account for many enterprises and businesses. The data has been stored and structured into usable format in order for analytics propose. However, some data has been storing in the format of hard copies, scanned document, images and PDFs which need to be transformed into digital form for future use. The objective of this paper is to propose the technique for recognizing the text from a physical document into digital format, by using Optical Character Recognition, also called OCR that attempts to extract all the text from photocopies into database structure. The experimental studies showed that the proposed technique makes the digitized documents completely searchable and editable with the average of accuracy performance around 75.38% for extracting attributes and 66.92% for extracting values from printed documents. This technology provides significant benefits to all businesses. Utilizing OCR helps businesses to easily seek for highly useful information throughout the document, and also there is a reduced amount of paper taking up space in the office.