Extracting topic maps from Web pages by Web link structure and content

Motohiro Mase, Seiji Yamada, Katsumi Nitta · 2008

We propose a framework to extract topic maps from a set of Web pages. We use the clustering method with the Web pages and extract the topic map prototypes. We introduced the following two points to the existing clustering method: The first is merging only the linked Web pages, thus extracting the underlying relationships between the topics. The second is introducing weighting based on the similarity from the contents of the Web pages and relevance between topics of pages. The relevance is based on the types of links with directories in the Web sites structure and the distance between the directories in which the pages are located. We generate the topic map prototypes by assuming that the clusters are the topics, the edges are the associations, and the Web pages related to the topics are the occurrences from the results of the clustering. Finally, users complete the prototype by labeling the topics and associations and removing the unnecessary items. We incrementally use a user’s evaluation of the topic maps to judge whether a Web page is unnecessary or necessary and then reduce the number of unnecessary pages. We use the relevance feedback along with a Support Vector Machine (SVM) to judge the Web pages. For this paper, at the first step, we mounted the proposed clustering method and conducted experiments to evaluate the effectiveness of extracting topic map prototypes. We eventually discussed the effectiveness of our two additional points by evaluating the extracted topic map prototypes.

Read the paper · More papers on PaperTik