Comparative analysis of big data management for social networking sites

Purti Beri, Sanjay Ojha · International Conference on Computing for Sustainable Global Development · 2016

This paper is focused on the topic “Big Data Management for Social Networking Sites”. Big data is a large term for larger data, which the traditional data processing techniques are unable to handle. In this paper review analysis of how big data is managed for social networking sites like Facebook and Twitter is done. The analysis of media content has been immense as Social Networking Sites are adding additionally massive chunk of the data and is not like other data sources where information produced and collected by SNS are not structured, also it is enormously large in volume that it is difficult to handle it by humans so such a large quantity of data which is not structured. For managing such huge amount of data we need Big Data Management. Here we analyze how Social Networking Sites like Facebook with over 1.4 billion active monthly users and Twitter with more than 500 million users, uses Hadoop Data Analysis Technologies like MapReduce, Pig and Hive for its big data management. At Facebook we see that, big stats on its system that processes nearly 2.5 billion contents and over five hundred terabytes of data every day. There are almost 2.7 billion user Like and around three hundred million photos daily loaded by users, which scans almost hundred and five terabytes of data every thirty minutes. For such enormous amount of data Facebook uses Hive, to store the data on HDFS the Hadoop distributed file system. On contrary Twitter has large data storage and for its processing, they have worked to implement a set of solutions storage inside Hadoop. It stores all the data which is LZO compressed, as this compression sustains a good balance between both speed and compression ratio in Hadoop.

Read the paper · More papers on PaperTik