An Architecture for Mining Massive Web Logs with Experiments

Andr As A. Bencz Ur, K. Aroly Csalog Any, Máté Uher · 2003

We introduce an experimental web log mining architecture with advanced storage and data mining components. The aim of the system is to give a flexible base for web usage mining of large scale Internet sites. We present experiments over logs of the largest Hungarian Web portal [origo] (www.origo.hu) that among others provides online news and magazines, community pages, software downloads, free email as well as a search engine. The portal has a size of over 130,000 pages and receives 6,500,000 HTML hits on a typical workday, producing 2.5 GB of raw server logs that remains a size of 400 MB per day even after cleaning and preprocessing, thus overrunning the storage space capabilities of typical commercial systems. As the results of our experiments over the [origo] site we present certain distributions related to the number of hits and sessions, some of which is somewhat surprising and different from an expected simple power law distribution. The problem of the too many and redundant frequent sequences of web log data is investigated. We describe our method to characterize user navigation patterns by comparing with choices of a memoryless Markov process. Finally we demonstrate the effectiveness of our clustering technique based on a mixture of singular value decomposition and refined heuristics.

Read the paper · More papers on PaperTik