Signature and Date-Based Document Image Retrieval
Ranju Mandal · Griffith Research Online · 2017
It is a common organisational practice nowadays to store and maintain large digital databases in an effort to move towards a paperless office. Large quantities of administrative documents are often scanned and archived as images (e.g. the ‘Tobacco’ dataset [1]) without adequate indexing information. Consequently, such practices have created a tremendous demand for robust ways to access and manipulate the information that such images contain. Manual processing (i.e. indexing, sorting or retrieval) of documents from these huge collections need substantial human effort and time. So, automatic processing of documents is required for office automation. In this context, Document Image Analysis (DIA) has enjoyed many decades of popularity as a research area to address these issues because of its huge application potential in many fields such as academics, banking and in industry. A document repository available for analysis in such a domain contains a large collection of heterogeneous documents. Automatic analysis of such large document database has been an interesting and challenging research field for many years, specifically due to the diverse layouts and contents. One way to efficiently search and retrieve documents from a large repository is to fully convert the documents to an editable representation (i.e. through Optical Character Recognition) and index them based on their content. There are many factors (e.g. high cost, low document quality, non-text components, etc.) which prohibit complete conversion of a document to an editable form. Hence, other components of a document, namely signatures, dates, logos, stamps/seals, etc. are worthy consideration for indexing, without the requirement for complete OCR.