Convolutional Autoencoder for Discriminating Handwriting Styles
Sanae Boutarfass, Bernard Besserer · 2019
We describe in this paper how we use a CNN-autoencoder for unsupervised anomaly detection, here adapted to detect anomalies in handwritten text samples. There is no distinction between a learning and test dataset, the text snippet containing the anomaly is the only data available. Since character recognition is not aimed at, we expect that the autoencoder should learn the styling characteristics of the writing. Algorithms used for graphology commonly measure the extends of the vertical strokes over or under the base line, the slope of the strokes or the closure of letters such as the “o”. Given the text baseline which is detected by image processing, the image of the sample is arbitrarily split in tiles for the training / recognition task. We also use a partial Radon projection to convert these tiles in a more abstracted representation for the learning/detection tasks. The Radon projection is computed over a [π /2] range which captures the most style variations in handwritten text. Discrimination takes place directly while training, and by shifting image tiles from batch to batch as the training epochs go on, outliers are discarded and localised.