Semicca: A new semi-supervised probabilistic CCA model for keyword spotting

Giorgos Sfikas, Basilis Gatos, Christophoros Nikou · 2017

In this paper we present a semi-supervised, attribute-based model suitable for keyword spotting (KWS) in document images. Our model can take advantage of available non-annotated segmented word images, as well as string annotations without a matching word image. We build our model by extending on the probabilistic interpretation of Canonical Correlation Analysis (CCA), solved using Expectation-Maximization (EM). On test-time, we back-project the query and database images to the embedded space by calculating the embedding space posterior density given the observations. Keyword spotting is then efficiently performed by computing query nearest neighbours in the embedded Euclidean space. We validate that our model offers superior performance given the presence of partially-labelled data, with keyword spotting trials on the Bentham and George Washington datasets.

Read the paper · More papers on PaperTik