Semi-automatic information retrieval and consolidation with a sample application

Petra Korica-Pehserl, Hermann A. Maurer · 2012

Imagine for a moment a Web where we can extract information from any website, know its context, and automatically assemble it with other information from other sources like databases, geo-maps, multimedia files etc. into a homogeneous document. Today the Web is populated by unstructured or semi-structured data, and the difficulty to consolidate it automatically makes this idea wishful thinking at the moment. This paper describes an attempt to use semi-automatic information retrieval and consolidation on the largest Austrian online encyclopedia Austria-Forum. We had access to the database of historic images of Austria with more than 40.000 images. We describe how we managed to incorporate some of those images suitable for the Austria-Forum without duplicates. The process comprised a large set of heuristics that accomplished a high percentage of the integration automatically. The results required only a moderate effort for a human expert to check if the images did indeed fit the entry proposed by the system.

Read the paper · More papers on PaperTik