Categorising AWS Common Crawl Dataset using MapReduce

Aatish Chiniah, Ayaz Chummun, Zaid Burkutally · 2019

Keeping track of websites connected to the Web is an impossible task given the amplitude and fluctuation of new sites being created and those going offline. In this paper we took the task to create a directory by categorising the websites using MapReduce. The dataset about the different websites are collected from the Amazon Web Service (AWS) Common Crawl dataset. Datasets are released at the end of each month and it contains information about all websites that are live during that month. After mining and translating the dataset, each URL are categorised based on its keywords and hosted country, the URLs are then placed in categories such as Art, Education, Shopping etc... After successfully assigning a category to each URL from the dataset, a new categorised dataset is updated with the new URL objects to be uploaded on the cloud. The data about each URL object is presented through a website whereby users are allowed to browse the data by both their categories and directory based domains. In order to evaluate the performance of the categorisation process, three techniques have been used which are Java 8 Streams, Multi-Threading and MapReduce. It has been observed that MapReduce was the most performing in terms of execution time, resource usage on Hadoop Cluster of seven nodes. 23576 (90.3%) URLs from an AWS datasets were successfully categorised and 5767(22.8%) non-English websites were successfully translated.

Read the paper · More papers on PaperTik