Content characterization using word shape tokens

Penelope Sibun, David Scott Farrar · 1994

By quickly classifying character images into character shape categories, it is possible to automatically extract syntactic information from the text of document images without optical character recognition. Using word shape tokens composed of these character shape codes, a properly trained text tagger can extract part-of-speech information from scanned document images. Later components of a document processing system can then use this information to locate topics, characterize document style, and assist in information retrieval.

Read the paper · More papers on PaperTik