A Bayesian Approach to Script Independent Multilingual Keyword Spotting
Gaurav Kumar, Venu Govindaraju · 2014
We propose a script independent Bayesian framework for keyword spotting in multilingual handwritten documents. The approach relies on local character level score and global word level hypothesis scores and learns a Bayesian logistic regression classifier to distinguish between keywords and non-keywords. In a Bayesian formulation of logistic regression, the integral over weights becomes intractable. Variational approximation is used for inference. In order to learn a robust classifier with minimal number of samples, we apply Bayesian active learning framework to request labels for those word images which provide maximum information gain in improving the classifier. We evaluate our system on multilingual datasets, publicly available IAM dataset for English, AMA for Arabic and LAW dataset for Devanagiri. The system is also evaluated on a synthetic multilingual dataset prepared by combining samples from IAM, AMA and LAW datasets. The results are comparable with the state of art multilingual keyword spotting framework.