Extracting Handwritten Regions In Japanese Document Images

Kim‐Ngan Nguyen, Thanh-Ha Do · 2020

Extracting handwritten regions in document images is an important task that has received a lot of interest in the image processing community because of its wide application in digitizing documents for recovery, storing, and searching purposes. Typically, the handwritten regions are found using different shapes features in document images such as the vertical and horizontal lines. Those approaches are time-consuming because the processing time depends largely on the complexity of the form (i.e., the number of information boxes). Moreover, we cannot use these features to gain more knowledge about the document. This paper presents a new approach to the problem that includes company's logo and form's title. First, each document is classified into different categories based on its logo and form's name and then the handwritten regions are extracted based on the type of each document. The classification stage is performed using a local feature descriptor. For performance comparison, four local descriptors (SIFT, SURF, BRIEF and ORB) were used to evaluate their performance in detecting and classifying logos in document image. A dataset of more than 2000 Japanese document images was built from 100 unpublished images with varying orientation and scaling to evaluate the accuracy of the proposed method.

Read the paper · More papers on PaperTik