Optimizing data partition for scaling out NoSQL cluster

Xiangdong Huang, Jianmin Wang, Yu Zhong, Shaoxu Song, Philip S. Yu · Concurrency and Computation Practice and Experience · 2015

Summary Data partition impacts the performance of Not Only SQL (NoSQL) systems significantly. Nowadays, many of the peer‐to‐peer NoSQL systems use consistent hashing to partition data automatically. These systems use virtual nodes and random data placement methods to divide the consistent hashing ring, which may lead to imbalanced data partition and degrade the overall system performance. The problem is prominent especially for scaling out heterogeneous clusters. Considering the capacity of each node, an imbalance coefficient of data distribution for a cluster is proposed firstly in this paper. Based on the imbalance coefficient, we propose a dynamic programming algorithm to calculate the position of the new coming node in the consistent hashing ring, which expands the consistent hashing ring more evenly without re‐shuffling the entire datasets. Simulations and experiments on Cassandra with Yahoo! Cloud Serving Benchmark (YCSB) benchmark show our algorithm is better than the state‐of‐the‐art work. Copyright © 2015 John Wiley & Sons, Ltd.

Read the paper · More papers on PaperTik