Three Experiments in Reliable Transaction Processing in RAID

Bharat Bhargava, Fady Lamaa, Pei-Jyun Leu, John Riedl · Purdue e-Pubs (Purdue University System) · 1988

This paper describes three studies that contribute towards the building of reliable distributed database systems.We measure the transaction processing times in the different phases of its execution in an operational database system called RAID.We find that response times of 300-500 milliseconds are achievable and much of it is due to the communications and commitment software.We have implemented and measured a concurrent checkpointing and rollback algorithm useful for dealing with failures of individual processes in the system.We find that the cost of synchronization for the coordinator process is of the order of the time required for taking a single checkpoint on the stable storage.We find that the concurrent execution does not reduce the message overhead or cpu usage but has the same communication delay as a single checkpoint instance.We show that much parallelism e.'cists in such algorithms.Finally we present two more experiments that were done by implementing a partially replicated database system.They measured the effects of the degree of replication and of the threshold representing the minimum number of copies.We find that a low degree of replication (25%) and a high threshold can provide the same data availability as a fully replicated database but with lower response time.

Read the paper · More papers on PaperTik