Text Retrieval from Document Images based on N-Gram Algorithm.

Chew Lim Tan, Sam Yuan Sung, Zhaohui Yu, Yi Xu · 2000

. We propose a method of text retrieval from document images using a similarity measure based on an N-Gram algorithm. We directly extract image features instead of using optical character recognition. Character image objects are extracted from document images based on connected components first and then an unsupervised classifier is used to classify these objects. All objects are encoded according to one unified class set and each document image is represented by one stream of object codes. Next, we retrieve N-Gram slices from these streams and build document vectors. Lastly, we obtain the pair-wise similarity of document images by means of the scalar product of the document vectors. Tests using four corpora of news articles have confirmed the validity of our method. 1 Introduction The Singapore National Library archives the entire set of past issues of major newspapers in Singapore in microfilms [1]. It is proposed that the microfilm images be digitized to facilitate retri...

Read the paper · More papers on PaperTik