Using Document Structure for Extracting Topics from the Web

Satoshi Oyama, Katsumi Tanaka · 2003

In this paper we propose a method for extracting keywords that detail the broad topic given by the user from the Web. Existing methods for extracting related keywords are based on term co-occurrence in each page and they extract many keywords irrelevant to the topic. Our method can precisely extract detailing topic keywords by considering positions (title or text) of keywords when they co-occur in documents. Using these keywords, the user can have a grasp of the topics in the Web and use them for more detailed search.

Read the paper · More papers on PaperTik