Handwriting pattern matching and retrieval with binary features

Bin Zhang, Sargur N. Srihari · 2003

This dissertation explores the role of binary features in pattern matching and retrieval with applications to handwriting processing. Specifically, this dissertation focuses on examining metric/non metric properties of similarity measures for binary feature vectors, accelerating pattern searching with nonmetric binary vector similarity measures, and developing effective binary features from handwritten word images for handwriting identification and retrieval tasks. Metric or non-metric verification of dissimilarity measures is important as efficient nearest neighbor (nn) search algorithms using metrics are usually inapplicable in searching a space based on non-metrics. We formulate a set of dissimilarity measures for binary vectors, and comprehensively examine their metric or non-metric properties. Especially, we discover and mathematically prove a special property, the Tri-Edge Inequality (TEI), associated with several measures. As the TEI property disables conventional fast search algorithms exploiting the triangle inequality, a cluster-tree based algorithm is developed to accelerate k-nn classification without any presuppositions about the metric form and properties of a dissimilarity measure. A mechanism of early decision making and minimal side-operations for choosing searching paths largely contribute to the efficiency of the algorithm. The algorithm is evaluated through extensive experiments. We develop a set of effective binary features from handwritten word images for tasks of handwriting identification, word image retrieval and signature authentication. Handwriting identification utilizes sixty-two alphanumeric characters and four characteristic words, from which binary features are extracted. Discriminability of each of sixty-two characters and four characteristic words is examined through two models of handwriting identification, writer identification and verification. Individuality of handwritten characters is quantitatively established. The ranking order of sixty-two alphanumeric characters in writer discriminability, for the first time, provides a scientific guidance for selecting most-discriminative characters in examining forensic documents. Binary features from sixty-two characters and four words lead to very high identification and verification rates. Word-level binary features are also applied to word image retrieval and signature authentication with very promising results. (Abstract shortened by UMI.)

Read the paper · More papers on PaperTik