Writer Verification using CNN Feature Extraction
Jun Chu, Mohammad Abuzar Shaikh, Mihir M. Chauhan, Lu Meng, Sargur N. Srihari · 2018
We propose an end-to-end learning method based on statistical features extracted on set-of-samples level as a step toward solving the writer verification problem which is about deciding whether two handwriting sources are identical given handwriting samples from the two sources. The set-of-samples features are extracted on top of single sample features. Single sample features are traditionally learned using sample-to-sample comparison similarity learning. In this paper, we learn it as a sub-module of an end-to-end two-sets-of-samples comparison similarity learning. We compare human-engineered (GSC features) single sample features and automatically learned features using convolutional neural networks (CNN) and find the latter performs better with abundant training data and data augmentation. The statistical features on sets of samples capture both inner-writer variabilities and intra-writer variabilities. Experiments are conducted on frequently occurring words and digraphs such as "and" and "th" from around 1500 writers. We perform experiments on pre-training using sample-to-sample similarity learning and end-to-end fine-tuning. The results show that two-sets-of-samples comparison gives much better accuracy than sample-to-sample comparison. In addition, the end-to-end training based on parametric statistical features gives better accuracy than standard distribution comparison tests such as the K-S test based on distance space.