1 On Analyzing Web Log Data: A Parallel Sequence Mining Algorithm
Ayhan Demiriz · 2003
Activities at enterprise-class web sites, as well as other web sites, are usually recorded via web logs. Collected logs consist of records from many click streams, which are defined as collections of hits (requests) from a specific user during a specific session. Using web logs is the most common way of collecting click stream data at this time. Thus data warehouses are built based on the crucial data extracted from web logs. This article proposes a parallel sequence mining algorithm, webSPADE, to analyze the click streams found in site web logs. In this process, raw web logs are first cleaned and inserted into a data warehouse. The click streams are then mined by webSPADE, the proposed algorithm, which uses one full scan and several partial scans of the data. An innovative web-based front-end is used for visualizing and querying the sequence mining results. By utilizing relational database technology, this analysis technique enables the analysis of very large amounts of data in a short amount of time.