Document Recognition for a Million Books

Sayeed Choudhury, Tim DiLauro, Robert Douglas Ferguson, Michael Droettboom, Ichiro Fujinaga · D-Lib Magazine · 2006

As initiatives such as Google Book Search (http://books.google.com/) and the Open Content Alliance (http://www.opencontentalliance.org/) advance efforts to digitize millions of books, there is great potential to make available vast amounts of information. To truly unlock this knowledge, however, it will be necessary to process the resulting digital page images to recognize important content, including both the semantic and structural aspects. Given the vast diversity of fonts, symbols, tables, languages and a host of other elements, it will be necessary to create flexible, modular, scalable document recognition systems. Document recognition involves extracting features from the images and even transcriptions of other documents in order to group diverse content.

Read the paper · More papers on PaperTik