Classification of Web Log Data to Identify Interested Users
Herit Trivedi, Narendra Limbad · International journal of advance research and innovative ideas in education · 2016
With the increasing demand of internet more number of website are used for getting required information and thus more usage of web-based data. Whereas the data that is stored in different type of format in form of web log file. This log file should be maintained as these data are in unsorted manner and it is done through preprocessing. Web usage mining focuses on discovering useful knowledge or information. Web log file is automatically generated by web server whenever user accesses the resource like webpage of website. Web Usage Mining consists of three steps, Data Preprocessing, Pattern Discovery and Pattern Analysis. Data Preprocessing extracts text format data form log file and store clean data into database. Pattern Discovery finds pattern, Classify data by applying mining techniques. Pattern analysis finds knowledge from the discovered pattern.The main objective of this thesis is instead of spending high amount of time in tracking the behaviour of overall users to redesign the web site, spend less amount of time in focusing interested group of users only.The existing model used Naive Bayesian Classification to identify interested group of users from web log data. In this we propose Classification based on Predictive Association Rules Mining (CPAR) algorithm to identify interested group of users and also we present a comparative study of Naive Bayesian with CPAR.