User preference-based data partitioning top-k skyline query processing algorithm
Zhiyun Zheng, Minghao Zhang, Mengyao Yu, Dun Li, Xingjin Zhang · 2021
To solve the problem of low efficiency of top-k skyline queries of massive data, a top-k skyline query algorithm based on user preference and data partitioning is proposed. First, according to the dimension priority given by the user, delete the dimension that the user is not interested in and reduce the original dataset. Second, in the MapReduce distributed environment, the dataset is divided into regions by combining user dimensional preferences and indifference thresholds to reduce the number of divided regions, and the local skyline datasets of each region are obtained through parallel calculations. Finally, the one-way dominance relationship between regions is used to gradually loosen control on the filtered dataset, obtain the global skyline dataset and output the first k data. Experimental results show that the query algorithm in this paper has obtained competitive results on open source synthetic datasets and real datasets compared with the baseline method.