Research on Fast and Parallel Clustering Method for Trajectory Data

Ne Wang, Shu Gao, Xiangwen Peng, Minrui Wang · 2018

In the era of big data, the development of satellite technology and Internet of Things has produced a large amount of trajectory data. We can effectively understand and predict the movement of the objects by analyzing their trajectory data. Now, most of density-based clustering algorithms have some disadvantages including the difficulty to determine input parameters, large I/O, and so on. DPC (Clustering by fast search and find of Density Peaks) is a new density-based clustering algorithm, which is simple and has only one input parameter, and also it is not affected by the data dimension, therefore, it can be effectively applied for trajectory clustering. However, in DPC, the local density is complex to calculate, and the cutoff distance is subjective to determine. In addition, DPC does not consider the existence of multiple cluster centers in the same cluster when clustering. To solve these problems, in this paper a fast clustering algorithm for trajectory data is put forward. In addition, Spark memory computing technology and data partitioning method are used to parallelize the algorithm, which greatly improves the clustering efficiency. Finally, experiments with three months' ship trajectory data from the Yangtze River have demonstrated that the clustering efficiency and effectiveness of our algorithm are significantly improved.

Read the paper · More papers on PaperTik