On an Algorithm for Identifying Sessions from Web Logs

Claudia Elena Dinucă, Dumitru Ciobanu · RePEc: Research Papers in Economics · 2011

The quality of decisions is based on the quality of processed data. So it is important that atthe beginning of the data mining process to providecorrect and quality data. The preprocessing datais a necessity for avoiding the failure of the dataanalysis. The idea that the data mining process canbe done without human supervision has proved to bewrong. Even so, the humans are trying toautomate as much as possible the process. From hereare resulting many algorithms and techniquesthat are implemented using various programming language. In this work is presented an algorithm foridentifying the sessions from a web logs file. It uses a value of 30 minutes to mark the end of asession and start another. We compute the average time for visiting the pages and using this we showthat the presented algorithm produces errors in identifying sessions. We consider that the correct wayto identify the session is to take into account theaverage time for visiting the pages.

Read the paper · More papers on PaperTik