RDF-Based Retrieval of Information Extracted from Web Product Catalogues

Ondřej Šváb, Martin Labský, Vojtěch Svátek · 2004

Extraction of relevant data from the raw source of HTML pages poses specific requirements on their subsequent RDF storage and retrieval. We describe an application of statistical information extraction technique (Hidden Markov Models) on product catalogues, followed with conversion of extracted data to RDF format and their structured retrieval. The domain-specific query interface, built on the top of Sesame repository, o#ers a simple form of navigational retrieval. Integration of further web-analysis methods, within the Rainbow architecture, is forthcoming.

Read the paper · More papers on PaperTik