Skewness and Nearest Neighbour Based Approach for Historical Document Classification
A Kavitha., Palaiahnakote Shivakumara, Govindaraju Hemantha Kumar · 2013
Classification of document is essential before feeding to OCR as there is no universal OCR which recognizes multiple scripts. Besides, classification of ancient historical documents such as Indus script is more challenging due to seal form inscribed on durable surfaces (stones) that does not have definite writing style. This result in characters may look different in different seals and non-uniform spacing between text lines. Therefore, in this paper, we propose two approaches, namely, Skew ness based Approach (SA) for Indus document classification from English and South Indian scripts and Nearest Neighbour based Approach (NNA) for classification of English from South Indian scripts. The SA explores the fact that skew ness between the components in the Indus document image with respect to x-axis is higher than skew ness between the components in English and South Indian documents. The NNA identifies the presence or absence of modifiers which are common in South Indian document images and are not present in English document images to study the straightness and cursive ness of the components for classification. The method is evaluated on 600 different document images, which include 100 documents of each type. The comparative study with existing methods shows that the proposed method is superior to existing methods in terms of classification rate.