RDF-Based Retrieval of Information Extracted from Web Product Catalogues
Ondřej Šváb, Martin Labský, Vojtěch Svátek · 2004
Extraction of relevant data from the raw source of HTML pages poses specific requirements on their subsequent RDF storage and retrieval. We describe an application of statistical information extraction technique (Hidden Markov Models) on product catalogues, followed with conversion of extracted data to RDF format and their structured retrieval. The domain-specific query interface, built on the top of Sesame repository, o#ers a simple form of navigational retrieval. Integration of further web-analysis methods, within the Rainbow architecture, is forthcoming.