Automatic content extraction of filled-form images based on clustering component block projection vectors
Hanchuan Peng, Xiaofeng He, Fuhui Long · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2003
Automatic understanding of document images is a hard problem. Here we consider a sub-problem, automatically extracting content from filled form images. Without pre-selected templates or sophisticated structural/semantic analysis, we propose a novel approach based on clustering the component-block-projection-vectors. By combining spectral clustering and minimal spanning tree clustering, we generate highly accurate clusters, from which the adaptive templates are constructed to extract the filled-in content. Our experiments show this approach is effective for a set of 1040 US IRS tax form images belonging to 208 types.