Large-Scale SMS Messages Mining Based on Map-Reduce
Tian Xia · 2008
Mining the popular SMS messages in a short period of time is very valuable. However, traditional OLAP-based mining method is not suitable for this very large scale dataset. In this paper, we present a mining approach based on Map-Reduce parallel framework: Firstly, original dataset is pre-processed and grouped by the senders' mobile numbers. Secondly, we do a transformation to regroup the dataset by the short content keys, and then extract the popular messages according to the count of different senders which have the same key. Furthermore, we propose a sentence similarity computation method and a novel Forward Merging and K-Neighbor Checking algorithm to merge the similar messages semantically. Experimental results show that the final dataset of popular messages is very small with high sending coverage ratio, and can meet the real requirements.