Towards Reverse Engineering of PDF Documents
Josef B. Baker, Alan P. Sexton, Volker Sorge · Czech digital mathematics library · 2011
Abstract. We present a progress report on our ongoing project of re-verse engineering scientific PDF documents. The aim is to obtain math-ematical markup that can be used as source for regenerating a document that resembles the original as closely as possible. This source can then be a basis for further document processing. Our current tool uses specialised PDF extraction together with image analysis to produce near perfect in-put for parsing mathematical formula. Applying a linear grammar and specific drivers for each output format to this input, we can produce an accurate reproduction of formulae when presented with their coordi-nates. In this paper we will show how this information can be exploited to discover the locations of both inline and display formulae, and also to perform rudimentary layout analysis of the whole document, identifying structures such as headings and paragraphs. 1