Information Extraction and Analysis on Certificates and Medical Receipts
Tzung‐Pei Hong, Wei-Chou Chen, Chih‐Hung Wu, Bo-Wen Xiao, Bing-Yang Chiang, Zhi-Xun Shen · 2022
Document digitalization has become a trend in recent years. It provides fast analysis and search because the information in the documents can be easily managed. However, in real applications, while digitalization is in progress, lots of information has not yet been digitalized and only stored on papers. A common demand of analyzing a large amount of documents would be a time-consuming mission because they need massive human labor. Nowadays, some computer vision algorithms have emerged and they can be applied in such a scenario. In this paper, we propose an automatic information extraction and analysis system for mandarin documents. It consists of three main steps. Firstly, the text regions in documents under natural scenes are detected. Secondly, these text regions are recognized and converted into digital forms. Finally, heuristic rules are designed and integrated into the system to improve the recognition accuracy. The proposed system is expected to eliminate the time-consuming problem of document information extraction.