Benchmarking the suitability of key-value stores for distributed scientific data
Hua Feng · 2012
This dissertation is going to explore the suitability for a key-value store to be utilized in a scientific environment. In the dissertation, we present a benchmark approach to evaluate the performance of a key-value store based on CAP (consistency, availability and partition tolerance) theorem, then explore the suitability when a scientific application needs to use a key-value store. With development of Internet technology, there is a trend that more and more data is generated for scientific research which leads to the demand of partitioning data across multiply nodes. However, the performance of relational database drops because of its data complexity. As a result, this dissertation will investigate the suitability of two key-value stores, Redis and Voldemort, by benchmarking it consistency and availability with different configuration (node number/dataset size). The result of benchmarking shows that Redis scales with the size of datasets but not number of nodes while Voldemort scales with the number of nodes but not size of datasets regarding consistency. The availability of Redis is better than Voldemort but both and Redis and Voldemort reduce its availability with more datasets. In addition, Voldemort shows a possibility that adding more nodes to Voldemort system may increase its availability. Furthermore, due to the client issues for both two data stores partition tolerance may not be evaluated currently. Finally, with an analysis of demanding scientific application, we conclude that both Voldemort and Redis could be suitable for a scientific application with dependence on the requirement of scientific application.