A Parallel Processing Technique for Extracting and Storing User Specified Data
Bannya Chanda, Shikharesh Majumdar · 2021
Users are often interested in a specific type of data (preferences) available from a large volume of data collected on the system. An efficient and effective system that can only store the user preferred data from the large raw data set helps the users to search for relevant information on time. The motivation behind this paper is to devise such a technique. The technique uses machine learning as well as parallel processing to efficiently filter out the user preferred data from a large raw data set. Firstly, the technique stores the filtered data and discards the remaining data which saves storage space for the user. Secondly, it leads to an enhanced searching experience for the user by reducing the search latency. Running the filtering operation can be CPU intensive which often leads to high latency for extracting user preferred data from the raw data set. To solve this problem, the technique employs parallel processing and machine learning, thus reducing the data filtering latency while making the data searching process faster for the user. A proof-of-concept prototype for this technique has been built on the Apache Spark parallel processing engine. The prototype is subjected to several performance experiments using synthetic datasets. The analysis of experimental results shows the viability of the proposed technique and provides insights into system behavior and performance.