RepliSmart

Ragini Krishna, Chandrasekhar Tekur, Arnab Phani · 2019

Replication of data in the context of databases is a way to improve the performance of queries (throughput). An ecosystem where data is replicated can also result in increased parallelism. With replicated data, there would be better fault tolerance. In some cases, replicating a set of data only in few nodes for higher efficiency (in terms of space), could be a choice. A particular set of data could be replicated in many nodes while others in only few, based on the access ratio of the data. Today, the decision of what data to be replicated on which all nodes, is taken based on few presumptions at the time of replication. Once the data is replicated, it remains in those nodes. Over a period of time, the requirements/queries accessing a set of data might change, and it may happen that the data that is less replicated might be the most desired, and vice versa.

Read the paper · More papers on PaperTik