Distorted Replicas: Intelligent Replication Schemes to Boost I/O Throughput in Document-Stores
Khaled Jouini · 2017
NoSQL databases commonly use an aggregate data model, where data that is expected to be accessed together is packed in a single clump and stored in a single node. This aggregate-orientation is essential to running on a cluster as it avoids cross-nodes joins and writes. An important downside of the aggregate data model is that it severely limits the ways data can be efficiently explored and processed, especially as NoSQL systems are being increasingly used by complex data-driven applications, mixing heterogeneous data access patterns.Replication is ubiquitous in NoSQL systems. In this work we propose a new replication scheme termed distorted replicas. Rather than being physically identical, distorted replicas are logically identical: they restructure replicated data in different ways, but keep the fundamental property of being constructible from one another. By doing so, distorted replicas provide new ways for exploring data, while still ensuring high availability. Experiments conducted in this paper show that even basic distortion schemes allow substantial performance improvements.