Towards mining past content of Web pages

Adam Jatowt, Katsumi Tanaka · New Review of Hypermedia and Multimedia · 2007

While much attention has recently focused on preserving the past content of the Web, there is still a lack of efficient tools for utilizing data stored in Web archives. Web archives constitute large data sources that could be extensively analysed and mined for knowledge discovery. In this paper, we describe the issues involved with mining Web archive data. We discuss several concepts related to collecting and analysing historical content of Web pages and briefly describe two knowledge discovery tasks—temporal summarization and object history detection.

Read the paper · More papers on PaperTik