Framework for Analyzing Web Access Logs using Hadoop and MapReduce
Pranjali Borgaonkar, Kumar Gaurav, Jyoti Yaduwanshi · 2018
Weblog mining is being a trending area of research which manages the extraction of interesting knowledge from logging data created by web servers. With the quick advancement of the web, web information mining exceptionally weblog mining plays a vital role in distinct field including customized data benefit enhancing plans, administrations of site et cetera. A few website specialists, web maintainers, website admins and web heads are unconscious of web server records, data each time they visit a site page. Essential logfile contains the data about username, IP address, Time Stamp Access request number, Number of bytes Exchanged, result status, error site pages, and URLs. Weblog documents can be utilized to give architect of site pages a substantially more detailed perspective of how clients are getting to their website. Weblog document can likewise be utilized to monitor usage patterns and employ them to enhance the outline and execution of sites. Moreover, tremendous weblogs can be utilized to investigate and upgrade an extensive variety of electronic things, for example, online instructional exercises, truth sheets, client records and exercises and so on by thinking about big data technologies. This digitally produced huge information from various sites is turning into a lot of information that the customary approach faces a bottleneck while processing such huge information In this paper, we propose a system for analysing web server logs of an educational organization through Hadoop and MapReduce which figures out the execution time.