Scalable Document Image Information Extraction with Application to Domain-Specific Analysis
Yingbin Zheng, Shuchen Kong, Wanshan Zhu, Hao Ye · 2019
Document images are ubiquitous, but existing methods mainly focus on the text reading but not information understanding. In this paper, we propose a novel document image information extraction framework with application to domain-specific analysis. Key gains of our system result from the modularized implementation of the document analysis modules needed for different document analysis problems. Further, we provide an efficient text recognition approach that makes a trade-off between performance and running speed for document images and a novel information extraction method with both visual and semantic information. Our framework is scalable and customizable, and only a few annotations of the keyword-content mapping is needed towards domain-specific document analysis.