A novel data preprocessing solution for large scale digital forensics investigation on big data
Heng Zhang · NORA - Norwegian Open Research Archives · 2013
ENGLISH: As the rapid development of high-technology, more and more novel and interesting applications and systems emerge. For example, people are willing to share their life any time any where just by accessing their Facebook accounts. In the same time, the popularity of mobility offices and fault-tolerance working platforms are becoming more and more hot than ever. For example, Dropbox is an popular cloud storage services among the world in recently. In addition, Google collaboration platform is one of the most successful business application for global users to work together in any time, even if they are not in the same office geologically. It is not difficult to find that more similar examples regarding to this concern. However, Nothing is prefect forever. Technology is a double-edged sword, especially in the information technology field. It pops up a lot of challenges. Consequently, digital forensics investigators pop up a significant question of how to implement large scale digital forensics investigation on big data effectively. It is impossible to handle those cases manually. However, some advanced techniques have been developed by research communities. For example, machine learning techniques are one of the most suitable candidate solutions to handle these big data cases. The significant merit for applying machine learning techniques is not only to introduce an automatic way of working, but also to process those complicated cases with higher precision than other means. Machine learning techniques consist of these stages, input gathering, data preprocessing, algorithm designing & deploying and output evaluation. The data preprocessing is an inevitable step for achieving better performance from machine learning techniques. However, research societies pay a lot of effort on advanced machine learning algorithm development and performance optimization. The crucial step of data preprocessing seems to be regarded by the same significance. This is the motivation for us to conduct this piece of work in this field. In this paper, we are going to address how to facilitate the implementation of largescale digital forensics investigation on big data set with the help of our data preprocessing solution. The methodology introduced in this paper is a hybrid solution based the stochastic theory, Grubbs’ criterion and the machine learning method, K Nearest Neighbour (KNN) algorithm. The complete technique contains two round of preprocessing work. While, the performance study on experiment results reflects a considerable achievement by our solution.