SIMILARITY MEASURES FOR WRITER CLUSTERING

J. Subrahmonia · 2004

This paper addresses the problem of improving the performance of an online, writer­independent, large­vocabulary, unconstrained, handwriting recognition sys­ tem by clustering writers with similar writing styles. Recognition performance is enhanced by identifying the writer cluster that a test writer is closest to and using a model trained for the corresponding writer cluster in decoding. The recognition system is based on hidden Markov models. A common set of features are computed for all writers, which are then projected to a lower dimensional space that preserves most of the information in the original feature set. The reduced dimensional space varies from writer to writer. This paper describes two measures of similarity between writing styles. The first is based on the distance between the writer­dependent reduced dimensional feature subspaces. The second is based on the hidden Markov Model output probabilities.

Read the paper · More papers on PaperTik