Duplication detection of a batch of image documents
Chen Hong-jia · Computer and Information Technology · 2003
Duplication detection of image documents based on database is an important content in the office automation.In this paper,a method is presented for detecting duplications of a batch of image documents based on cluster analysis.First,converts a page of document have read into computer to binary bitmap.Giving a series of interlocking concentric disk(The center of all disks is computed according to the edge of this page),computing radial pixel densities (the number of ‘on’ pixels in each annuli)as the feature vector.Establishing a distance among feature vectors, and detecting duplications by cluster analysis.The result of stimulating experiments by MATLAB,85%~98% of the documents got from the internet can be classify correctly.