Automatic information extraction from the web

Neva Smith, King-Ip Lin · 2012

We implemented a system that automatically extract information from web pages and store them in a database system for easy processing. Our system only requires to user to provide the underlying database structure to work. In addition, our system also make use of existing information we extracted to improve accuracy. We use the recipe domain to illustrate our work, and initial experiment results (of 70% accuracy) provide encouraging signs that this method will be effective on other domains.

Read the paper · More papers on PaperTik