Clustering Social Networking Data With K-Means Algorithm Using R Language

Sujeet Kumar Sahani, Sonam Singh · International Journal of Scientific Research in Computer Science Engineering and Information Technology · 2024

The main objectives of this research work are to report detailed empirical studies on sequential and parallel algorithms for diverse clustering tasks executed on very large social network datasets using memory efficient out-of-core approaches. We evaluate the spark implementation for R on Cloudera using the data from social media review datasets like k-means and hierarchical clustering to rank these algorithms. This implementation leverages the YouTube dataset from UCI Machine Learning Repository. Our goal is to compare a few algorithms, so we can know exactly how accurately these models are performing. Ultimately we want to deal with testing and ranking clustering method, and mining and finally clustering massive amounts of unstructured data.

Read the paper · More papers on PaperTik