Collaborative Information Extraction and Mining from Multiple Web Documents

Tak-Lam Wong, Wai Chun Lam, Shing-Kit Chan · 2006

We develop an unsupervised framework which can collaboratively extract information from multiple Web pages, as well as conduct feature mining tasks in a unified model. Our model allows tight interactions of the two tasks removing the unnecessary boundary between the two tasks. It is beneficial for both tasks since the decisions for information extraction and feature mining can be done in a coherent manner assigning solutions optimizing the quality of both tasks and at the same time eliminating the potential conflicts. Our approach is designed based on an undirected graphical model which can model the inter-dependence between the neighbouring tokens within the same Web page, as well as tokens in different Web pages. Multiple Web pages are considered under this model and the information can be extracted collectively. This design also leads to another characteristic of our framework in that it can conduct mining across Web pages simultaneously. We demonstrate the efficacy of our model by applying it to the important product feature mining application. Extensive experiments on real-world data have been conducted to evaluate our framework.

Read the paper · More papers on PaperTik