AUTOMATIC DOCUMENT STRUCTURE ANALYSIS OF STRUCTURED PDF FILES

Rosmayati Mohemad, Zulaiha Ali Othman Abdul Razak Hamdan · International Journal of New Computer Architectures and their Applications · 2011

Nowadays, government agencies, business corporate companies and personal individual disseminate their electronic documents in Portable Document Format (PDF) as the quickest and convenient way to publish information including text, tables and graphical images. In the survey conducted by Association of Information and Image Management (AIIM) in 2008, 90% of 200 member organizations stored information in PDF either by scanning the documents or converting from Microsoft Office files to PDFbased format and it is predicted the use of PDF fluctuates to 93% for the next five years [1]. Various types of digital documents either newspapers, catalogues, reports, magazines, articles and even forms are available in PDF since the technology is good at offering open standard feature in which it is platform independent for sharing, archiving, retrieving and printing electronic document. Despite of it benefits, PDF has drawback in terms of content and structure analysis. As the result, information represented in PDF format is inconvenient for people to retrieve and reuse in other applications automatically, for example decisionmaking applications and office automation systems, which requires International Journal on New Computer Architectures and Their Applications (IJNCAA) 1(2): 404-411 The Society of Digital Information and Wireless Communications, 2011 (ISSN: 2220-9085)

Read the paper · More papers on PaperTik