Text Content Based Layout Analysis
José Ramón Fernández Prieto, Vicente Bosch, Enrique Vidal, Dominique Stutzmann, Sébastien Hamel · 2020
State-of-the-art Document Layout Analysis methods rely on graphical appearance features in order to detect and classify the different layout regions present in a scanned text image. In many cases, however, performing this task using only graphical information is problematic or impossible. Only by actually reading some text in the boundaries of the problematic regions it becomes possible to reliably detect and separate these regions. In these situations, textual, content-based features would be required, but since transcription is usually performed after layout analysis, a vicious circle arises. In this work, we circumvent this deadlock by making use of the recently introduced concept of Probabilistic Index Map. We use the word relevance probabilities provided by this map to calculate relevant text content based features at the pixel level. We assess the impact of these new features on a historical document complex paragraph classification task. The experiments are performed using both a classical Hidden Markov Model approach and Deep Neural Networks. The obtained results are encouraging and showcase the positive impact text content based features will have on the Document Layout Analysis research field.