A Novel Algorithm Using Content-Based Filtering Technology in Apache Spark for Big Data Analysis

Yolamu Kamukwamba, Liu Chunxiao · 2021

Big data has been one of the fastest-growing research areas that many researchers are into, with large amounts of data being uploaded to the internet every minute analyzing this data can benefit businesses and content creators. Big companies are already making good use of customer data and can predict what customers might need and want based on their recent activities. In this paper, I am going to propose a novel algorithm that makes the process of analyzing big data much faster and easier by using content-based filtering on structured data in Apache Spark. We used models such as the classification model to classify the data in relevant categories that we need and find the relationship between days it takes to trend with views then use a predictive model to predict what type of content is good to produce. We used this on an existing realworld dataset and we were able to get good results. The results are encouraging and it proves that this is an improvement over traditional dig data analysis methods that uses unstructured data and deep learning for big data analysis when working on a very big dataset.

Read the paper · More papers on PaperTik