Handwritten Gujarati Word Image Matching using Autoencode
Kathiriya B. Khushali, Mukesh M. Goswami, Suman Kumar Mitra · 2020
Gujarati is the native language of the state of Gujarat. In worldwide more than 65.5 million people speak Gujarati. Large numbers of Gujarati printed and penmanship documents are accessible in digital format. However, fetching information in these digital document images is a crucial yet important task. Proposed work focuses on Gujarati handwritten word-matching. As it is very difficult to extract features manually from handwritten word images, a convolutional Autoencoder model is deployed to automatically learn optimum features from data. The trained model is expected to convert the two-dimensional word-image into a compact feature vector without losing the discriminative strength of features. A moderate-sized Gujarati handwritten word-image database was also generated for experiment purpose since there is no such database available in the literature. The initial results are encouraging and warrant further investigation in this direction.