Predicting Semantic Labels of Text Regions in Heterogeneous Document Images

Somtochukwu Enendu, Johannes C. Scholtes, Jeroen B. J. Smeets, Djoerd Hiemstra, Mariët Theune · University of Twente Research Information · 2019

This paper describes the use of sequence labeling methods in predicting the semantic labels of extracted text regions of heterogeneous electronic documents, by utilizing features related to each semantic label. In this study, we construct a novel dataset consisting of real world documents from multiple domains. We test the performance of the methods on the dataset and offer a novel investigation into the influence of textual features on performance across multiple domains. The results of the experiments show that the neural net-work method slightly outperforms the Conditional Random Field method with limited training data available. Regarding generalizability, our experiments show that the inclusion of textual features aids performance improvements.

Read the paper · More papers on PaperTik