Inferring Structure and Meaning of Semi-Structured Documents by using a Gibbs Sampling Based Approach

Vinodh Kumar Ravindranath, Devashish Deshpande, K Venkata Vijay Girish, Darshan Patel, Neel Jambhekar, Vikash Singh · 2019

In this paper, we present a novel and elegant method to extract the structure and derive meaning from semi-structured text documents such as resumes. Semi-structured text documents have an information hierarchy captured by use of textual styling, formatting and visual layout organization. So there is a need to build an algorithm that can both extract the structural hierarchy as well as classify the text blocks into known categories. Our algorithm proposes a generative statistical model, that generates the corpus of semi-structured documents, in a way that allows us to apply the Gibbs sampling algorithm to estimate the parameters of the underlying model that results in extraction of both structure and meaning from the documents. We have seen that applying our Gibbs sampling based algorithm on a random set of 1000 resumes has an extraction accuracy of close to 75%.

Read the paper · More papers on PaperTik