Content-Based Textual Big Data Analysis and Compression
Fei Gao, Ananya Dutta, Jiangjiang Liu · 2018
With the growing enhancement of technology and the Internet, the number of people who are using the Internet is increasing daily. Users are engaged in web searching and accessing different types of websites, such as social media, banking, etc. As a result, a large volume of data is being generated in every day. It is necessary to load this data for analysis purposes. However, memory space and transmission time are the most important factors of limited processing. In most cases, we only need to extract the important textual data from these vast raw datasets. In this work, we propose content-based compression (CBC) for textual data analysis on the basis of the Huffman Code. The data is pre-analyzed to find very high frequency words and then a shorter symbol is inserted to replace those words. This compression approach is performed in an effort to maintain the original format of the data so that, compressed data structure could be completely transparent to Hadoop platform. The algorithm is evaluated on a set of real world data sets (e.g. Amazon movie review, food review, etc.) and a 52.4% average data size reduction is obtained from the experiment. Though this gain may seem modest, this can be supplementary to all other compression optimization techniques. Furthermore, the proposed technique can be effectively applied for the big data optimization purpose.