A Co-training based Framework for Writer Identification in Offline Handwriting.
Utkarsh Porwal, Venu Govindaraju · 2011
Abstract—Traditional forensic document analysis methods have focused on feature-classification paradigm where a machine learning based classifier is used to learn discrimination among multiple writers. However, usage of such techniques is restricted to availability of a large labeled dataset which is not always feasible. In this paper, we propose a Cotraining based approach that overcomes this limitation by exploiting independence between multiple views (features) of data. Two learners are initially trained on different views of a smaller labeled training data and their initial hypothesis is used to predict labels on larger unlabeled dataset. Confident predictions from each learner are used to add such data points back to the training data with predicted label as the ground truth label, thereby effectively increasing the size of labeled dataset and improving the overall classification performance. We conduct experiments on publicly available IAM dataset and illustrate the efficacy of proposed approach.