Synchronizing Small Data in a big world
Keith Barber, Khaled Harfoush · 2014
Distributed databases have received a tremendous amount of attention in recent years due to the explosive growth in data facilitated by the Internet. This trend has further accelerated with the emergence of easy to use “Big Data” processing techniques such as Map/Reduce in combination with infrastructure and platform as a service offerings such as AmazonWeb Services. Distributed databases have tended to focus on partitioning and replicating data across a cluster, since that is the most common inhibitor to scalability. However, there are cases where the data is relatively bounded and has higher availability requirements such as very low latency, which has not been so thoroughly studied. In this paper we investigate the scalability properties of “Small Data”, in which a large set of small values needs to be globally replicated with latency dominating consistency requirements.